跳到正文
Hacker News · AI· amoursy·· 3 小时前AI 评分58

HieraticBench:13 个模型能否读懂古埃及僧侣体手写

Show HN: HieraticBench – Can AI read ancient Egyptian handwriting?

AI 导读

Veeza AI 创始人 Aly Moursy 发布 HieraticBench,用一位牛津埃及学家 2022 年手写的一句僧侣体作为密封考题,测试 13 个模型能否识别并翻译这种古埃及日常手写体。

正文
A single line of hieratic handwriting in black ink, about a dozen joined signs, written in 2022 by an Oxford Egyptology professor.
  • Claude Opus 5.5 said Not a known writing system. Wrong.
  • Claude Fable 5.1 said Not a real writing system. Wrong.
  • Claude Fable 5.1 said Hieratic, Demotic, Meroitic, Pahlavi or Elbasan. None confirmed. Wrong.
  • Claude Fable 5.1 said Amenemhat, “Amun is in front”. Wrong.
  • Claude Haiku 4.5 said Urdu, “One should learn every day”. Wrong.
  • Claude Fable 5.1 said Pahawh Hmong. Wrong.
  • Claude Fable 5.1 said Kharosthi. Wrong.
  • Claude Fable 5.1 said Old Uyghur. Wrong.
  • Claude Sonnet 5.5 said Nüshu. Wrong.
  • GPT-6 Astra said Tibetan. Wrong.
  • GPT-6 Astra said Geba. Wrong.
  • GPT-6 Astra said Marchung. Wrong.
  • GPT-6 Astra said Marchen. Wrong.
  • GPT-6.1 Sol said Tibetan. Wrong.
  • GPT-6.1 Sol said Han. Wrong.
  • GPT-6.1 Sol said Enochian. Wrong.
  • Gemini 3.1 Pro said Paleo-Hebrew. Wrong.
  • Gemini 3.1 Pro said Reformed Egyptian. Wrong.
  • Gemini 3.8 Flash said Nabataean. Wrong.
  • Gemini 3.8 Flash said Linear Elamite. Wrong.
  • Gemini 3.8 Flash said Proto-Sinaitic. Wrong.
  • Gemini 3.8 Flash said Tibetan. Wrong.
  • Grok 4.7 said Arabic. Wrong.
  • Grok 4.7 said Georgian Mkhedruli. Wrong.
  • Grok 4.7 said Baybayin. Wrong.
  • Kimi K3 said Latin. Wrong.
  • Kimi K3 said Kharosthi. Wrong.
  • Llama 4 Maverick said Mongolian. Wrong.
  • Llama 4 Maverick said Hangul. Wrong.
  • Llama 4 Maverick said Mkhedruli Georgian. Wrong.
  • Mistral Medium 3.5 said Rongorongo. Wrong.
  • Mistral Medium 3.5 said SINHALA. Wrong.
  • Mistral Medium 3.5 said Thai. Wrong.
  • Mistral Medium 3.5 said Syriac. Wrong.
  • Qwen3.8 Max said Mongolian script. Wrong.

An Oxford Egyptologist wrote it for me in 2022, in hieratic, the everyday handwriting of ancient Egypt. We tested 13 models. Not one could even name the script.

Aly Moursy, Founder of Veeza AI (YC F26)

What the models said.

The same image, the same plain question. Asked in the apps people actually use, and through the API.

  1. Claude Opus 5.5

    4 October 2026, Claude app

    Can you identify and translate this script?

    I can't confidently identify this as any known writing system, so I can't give you a translation. I'd rather say so than invent one.

    It doesn't match Chinese, Japanese, Korean, Tangut, Khitan, Jurchen, Nüshu, or the common shorthand systems like Gregg or Pitman.

    Did not recognise hieratic

  2. Claude Fable 5.1

    September 2026, Claude app

    Can you translate this script? First identify what it is and then translate it!

    I looked at this carefully, and I don't think it's a real writing system, so I can't give you a translation without making one up.

    Did not recognise hieratic

  3. Claude Fable 5.1

    September 2026, Claude app

    Follow-up after the cold attempt: This may be a script that's no longer in use or a reconstruction of one.

    Even against extinct and reconstructed scripts, I still can't place it, so I'd rather not guess at a translation.

    It listed Egyptian hieratic and Demotic among the closest visual families, then rejected them.

    Named the right family only after a nudge, then rejected it

  4. Claude Fable 5.1

    September 2026, Claude app

    Told the image was hieratic and that a speculative transcription was welcome.

    What survives narrowing: a personal name built on Amun, most likely Amenemhat.

    Confident sign-by-sign reading. Wrong.

  5. Claude Haiku 4.5

    5 October 2026, API, HieraticBench harness

    Can you identify and translate this script?

    This appears to be Urdu script (a variant of Persian/Arabic script).

    It then supplied a word-by-word Urdu reading and translated it as “One should learn every day.”

    Wrong script, invented translation

Through the API

Both images, several tries each, the same plain question

Claude Fable 5.1
Unknown 3 times, Pahawh Hmong once, Kharosthi once, Old Uyghur once.
Claude Haiku 4.5
Urdu 4 times, Unknown twice.
Claude Opus 5.5
Unknown 6 times.
Claude Sonnet 5.5
Unknown 5 times, Nüshu once.
GPT-6 Astra
Tibetan twice, Geba once, Marchung once, Unknown once, Marchen once.
GPT-6.1 Sol
Tibetan 3 times, Han once, Enochian once, Unknown once.
Gemini 3.1 Pro
Unknown 3 times, Paleo-Hebrew twice, Reformed Egyptian once.
Gemini 3.8 Flash
Nabataean twice, Linear Elamite once, Proto-Sinaitic once, Unknown once, Tibetan once.
Grok 4.7
Arabic twice, No answer once, Unknown once, Georgian Mkhedruli once, Baybayin once.
Kimi K3
Unknown 4 times, Latin once, Kharosthi once.
Llama 4 Maverick
Mongolian 3 times, Hangul twice, Mkhedruli Georgian once.
Mistral Medium 3.5
Rongorongo once, Unknown once, SINHALA once, Sinhala once, Thai once, Syriac once.
Qwen3.8 Max
Unknown 5 times, Mongolian script once.

Opus 5.5 checked Tangut, Khitan, Jurchen and Nüshu. It checked Gregg shorthand. It never checked Egypt.

Astra also failed to name the script. We didn't keep the transcript.

They know what a papyrus looks like. They can't read the writing.

Show the best model a real hieratic document, a papyrus, a pottery shard or a plate from an Egyptology book, and it names the script 95% of the time. Show it one sentence in the same script, written fresh by an expert, and it scores 0% across 78 tries.

Ask it which hieroglyph a single hieratic sign stands for and the best model is right 13% of the time. Models seem to know what a page of hieratic looks like. They don't know the signs it is made of.

The goal isn't this sentence. It's hieratic.

The sentence is the exam. Its answer has never been published and isn't stored anywhere, so a model can't pass by remembering something it read online. The only way through is to actually read hieratic.

What we want is AI that can read the handwriting of ancient Egypt, sign by sign, on any papyrus or pottery shard, and help the few people who read it today with the many documents still waiting. A model that can do that will read this sentence too.

Hieroglyphs were for stone. Hieratic was for everything else.

Hieratic is the cursive form of Egyptian hieroglyphs, written fast with a reed brush on papyrus, pottery shards and wooden boards. For more than three thousand years it was how Egypt actually wrote. Letters, tax records, medical manuals, maths problems, love poems, the stories people told.

The Rhind Mathematical Papyrus is hieratic. So is the Edwin Smith surgical papyrus, the Tale of Sinuhe, and the Diary of Merer, a logbook kept during the building of the Great Pyramid and the oldest inscribed papyrus ever found.

Hieroglyphs get the museum walls. Hieratic is where the people are. And almost nobody alive can read it.

Prisse Papyrus, Teaching of Ptahhotep (BnF Égyptien 190), written in hieratic.

Prisse Papyrus, Teaching of Ptahhotep (BnF Égyptien 190), c. 1800 BCE. Scribes wrote headings in red ochre. The red on this site comes from them. Zunkir, via Wikimedia Commons. CC BY-SA 4.0.

Why this is a fair test.

The answer is sealed.
The translation has never been published and the benchmark doesn't store it anywhere, so no model can have memorised it.
Almost nobody can read it.
Hieratic is read fluently by a small circle of specialists, so there is very little labelled data to learn from. Reading it means learning, not recall.
Every scribe wrote differently.
Signs fuse into ligatures and drift across centuries and hands. A model has to generalise from a sign list to a living hand.
Progress shows up early.
Naming the script and reading single signs are scored automatically on public data. Partial credit is real credit.

Four rungs, from seeing to reading.

Every rung after the first tells the model it is looking at hieratic, so a model that can't name the script still gets a fair shot at reading it.

  1. 1

    Identify

    What writing system is this?

    Asked cold, the way you would ask a friend. Demotic and hieroglyphic controls sit in the set, so answering hieratic every time doesn't pay.

    Best so far 95% on real documents, 0% on the sentence

  2. 2

    Signs

    Which hieroglyph is each sign?

    Answered in Gardiner sign-list codes, the standard Egyptologists use. Scored by sign error rate.

    Best so far 13%

  3. 3

    Transliterate

    How does it sound?

    Standard Egyptological transliteration. Not scored on the sentence, since no answer is stored.

    Not scored yet

  4. 4

    Translate

    What does it say?

    English. Not scored on the sentence either. Scoring is built in for future public texts.

    Not scored yet

Leaderboard

Every score comes from the open harness, through Anthropic's API for Claude and OpenRouter for everyone else. The first two columns ask a model to name the script, on the professor's sentence and on real ancient documents. Updated 5 October 2026.

How scoring works

  1. 1

    Claude Opus 5.5 (high)

    Sentence
    0%
    Documents
    95%
    Signs
    13%
  2. 2

    Claude Fable 5.1 (high)

    Sentence
    0%
    Documents
    95%
    Signs
    8.7%
  3. 3

    Claude Sonnet 5.5 (high)

    Sentence
    0%
    Documents
    94%
    Signs
    11%
  4. 4

    Kimi K3 (high)

    Sentence
    0%
    Documents
    84%
    Signs
    Not run
  5. 5

    Qwen3.8 Max (high)

    Sentence
    0%
    Documents
    81%53 of 87
    Signs
    Not run
  6. 6

    GPT-6.1 Sol (high)

    Sentence
    0%
    Documents
    79%
    Signs
    Not run
  7. 7

    Llama 4 Maverick

    Sentence
    0%
    Documents
    61%
    Signs
    Not run
  8. 8

    Mistral Medium 3.5 (high)

    Sentence
    0%
    Documents
    46%
    Signs
    Not run
  9. 9

    Claude Haiku 4.5

    Sentence
    0%
    Documents
    43%
    Signs
    0.7%
  10. 10

    Grok 4.7 (high)

    Sentence
    0%
    Documents
    25%
    Signs
    Not run
  11. 11

    GPT-6 Astra (high)

    Sentence
    0%
    Documents
    Not run
    Signs
    Not run
  12. 12

    Gemini 3.1 Pro (high)

    Sentence
    0%
    Documents
    Not run
    Signs
    Not run
  13. 13

    Gemini 3.8 Flash (high)

    Sentence
    0%
    Documents
    Not run
    Signs
    Not run
RankModelThe sentenceReal documentsSingle signs
1Claude Opus 5.5 (high)0%95%13%
2Claude Fable 5.1 (high)0%95%8.7%
3Claude Sonnet 5.5 (high)0%94%11%
4Kimi K3 (high)0%84%Not run
5Qwen3.8 Max (high)0%81%53 of 87Not run
6GPT-6.1 Sol (high)0%79%Not run
7Llama 4 Maverick0%61%Not run
8Mistral Medium 3.5 (high)0%46%Not run
9Claude Haiku 4.50%43%0.7%
10Grok 4.7 (high)0%25%Not run
11GPT-6 Astra (high)0%Not runNot run
12Gemini 3.1 Pro (high)0%Not runNot run
13Gemini 3.8 Flash (high)0%Not runNot run

Claude Fable 5.1 (high) declined 6 questions, which count as wrong. Real documents counts hieratic documents only, not the demotic and hieroglyphic controls. Transliterating and translating the sentence aren't scored, because no answer is stored anywhere. Chat-app transcripts above are quoted for the record and never counted here.

The dataset

268 items. 87 real hieratic documents, 29 controls in other Egyptian scripts, 150 single signs and 2 sealed sentences. Every public image is openly licensed and credited.

Browse every item

The professor's sentence

2022

Sealed© Aly Moursy. Free to use for evaluating models.

The scribe's rendition

2022

Sealed© Aly Moursy. Free to use for evaluating models.

Papyrus Berlin 3022 (Story of Sinuhe), opening section

Ägyptisches Museum und Papyrussammlung, Berlin

HieraticCC BY-SA 4.0

Ostracon with Pharaoh Spearing a Lion and a Royal Hymn on its Back, Met 26.7.1453

ca. 1200–1080 BCE, The Metropolitan Museum of Art, New York

HieraticCC0 (Met Open Access)

Rhind Mathematical Papyrus (BM EA 10057), detail

c. 1650 BCE, British Museum, London (EA 10057)

HieraticPublic domain

Ostracon: letter from the scribe Mose to the vizier Pesiur (Ashmolean HO 71)

13th century BCE, Ashmolean Museum, Oxford (HO 71)

HieraticCC BY-SA 4.0

Marriage Contract, Met 35.4.1a, b

380–343 BCE, The Metropolitan Museum of Art, New York

DemoticCC0 (Met Open Access)

Single hieratic sign, London, British Museum, EA 9999 (AKU HT 10440)

New Kingdom, Dynasty 20, Ramesses III, regnal year 32, month 3, season šm.w, day 6

Single signCC BY 4.0

Help AI learn to read hieratic.

This only works with more people. It needs the few who can read hieratic, and the people building the models that one day might.

Egyptologists
Write us a sentence. Every new sealed sentence makes the benchmark harder to game and the result harder to argue with. Sign annotations on published papyri help just as much.
Offer a sentence
AI researchers
Run the harness on your model with your own key, or build a reader from scratch. Every result goes on the public leaderboard.
Run the benchmark
Everyone else
Share it. Introduce us to an Egyptologist. Sponsor a commissioned sentence. The benchmark grows one sentence at a time.
Get in touch

Or go see it for yourself.

Egypt's museums are full of hieratic, on papyrus and on pottery shards. Deir el-Medina, near Luxor, is the village where the workmen who built the royal tombs left thousands of notes in it.

Need a visa? Veeza AI handles the Egypt e-visa before you fly.

Get your Egypt e-visa

The Nile
Abu Simbel

来源:Hacker News · AI · hieraticbench.vercel.app