HieraticBench:13 个模型能否读懂古埃及僧侣体手写
Show HN: HieraticBench – Can AI read ancient Egyptian handwriting?
Veeza AI 创始人 Aly Moursy 发布 HieraticBench,用一位牛津埃及学家 2022 年手写的一句僧侣体作为密封考题,测试 13 个模型能否识别并翻译这种古埃及日常手写体。

- Claude Opus 5.5 said Not a known writing system. Wrong.
- Claude Fable 5.1 said Not a real writing system. Wrong.
- Claude Fable 5.1 said Hieratic, Demotic, Meroitic, Pahlavi or Elbasan. None confirmed. Wrong.
- Claude Fable 5.1 said Amenemhat, “Amun is in front”. Wrong.
- Claude Haiku 4.5 said Urdu, “One should learn every day”. Wrong.
- Claude Fable 5.1 said Pahawh Hmong. Wrong.
- Claude Fable 5.1 said Kharosthi. Wrong.
- Claude Fable 5.1 said Old Uyghur. Wrong.
- Claude Sonnet 5.5 said Nüshu. Wrong.
- GPT-6 Astra said Tibetan. Wrong.
- GPT-6 Astra said Geba. Wrong.
- GPT-6 Astra said Marchung. Wrong.
- GPT-6 Astra said Marchen. Wrong.
- GPT-6.1 Sol said Tibetan. Wrong.
- GPT-6.1 Sol said Han. Wrong.
- GPT-6.1 Sol said Enochian. Wrong.
- Gemini 3.1 Pro said Paleo-Hebrew. Wrong.
- Gemini 3.1 Pro said Reformed Egyptian. Wrong.
- Gemini 3.8 Flash said Nabataean. Wrong.
- Gemini 3.8 Flash said Linear Elamite. Wrong.
- Gemini 3.8 Flash said Proto-Sinaitic. Wrong.
- Gemini 3.8 Flash said Tibetan. Wrong.
- Grok 4.7 said Arabic. Wrong.
- Grok 4.7 said Georgian Mkhedruli. Wrong.
- Grok 4.7 said Baybayin. Wrong.
- Kimi K3 said Latin. Wrong.
- Kimi K3 said Kharosthi. Wrong.
- Llama 4 Maverick said Mongolian. Wrong.
- Llama 4 Maverick said Hangul. Wrong.
- Llama 4 Maverick said Mkhedruli Georgian. Wrong.
- Mistral Medium 3.5 said Rongorongo. Wrong.
- Mistral Medium 3.5 said SINHALA. Wrong.
- Mistral Medium 3.5 said Thai. Wrong.
- Mistral Medium 3.5 said Syriac. Wrong.
- Qwen3.8 Max said Mongolian script. Wrong.
An Oxford Egyptologist wrote it for me in 2022, in hieratic, the everyday handwriting of ancient Egypt. We tested 13 models. Not one could even name the script.
Aly Moursy, Founder of Veeza AI (YC F26)
What the models said.
The same image, the same plain question. Asked in the apps people actually use, and through the API.
Claude Opus 5.5
4 October 2026, Claude app
Can you identify and translate this script?
I can't confidently identify this as any known writing system, so I can't give you a translation. I'd rather say so than invent one.
It doesn't match Chinese, Japanese, Korean, Tangut, Khitan, Jurchen, Nüshu, or the common shorthand systems like Gregg or Pitman.
Did not recognise hieratic
Claude Fable 5.1
September 2026, Claude app
Can you translate this script? First identify what it is and then translate it!
I looked at this carefully, and I don't think it's a real writing system, so I can't give you a translation without making one up.
Did not recognise hieratic
Claude Fable 5.1
September 2026, Claude app
Follow-up after the cold attempt: This may be a script that's no longer in use or a reconstruction of one.
Even against extinct and reconstructed scripts, I still can't place it, so I'd rather not guess at a translation.
It listed Egyptian hieratic and Demotic among the closest visual families, then rejected them.
Named the right family only after a nudge, then rejected it
Claude Fable 5.1
September 2026, Claude app
Told the image was hieratic and that a speculative transcription was welcome.
What survives narrowing: a personal name built on Amun, most likely Amenemhat.
Confident sign-by-sign reading. Wrong.
Claude Haiku 4.5
5 October 2026, API, HieraticBench harness
Can you identify and translate this script?
This appears to be Urdu script (a variant of Persian/Arabic script).
It then supplied a word-by-word Urdu reading and translated it as “One should learn every day.”
Wrong script, invented translation
Through the API
Both images, several tries each, the same plain question
- Claude Fable 5.1
- Unknown 3 times, Pahawh Hmong once, Kharosthi once, Old Uyghur once.
- Claude Haiku 4.5
- Urdu 4 times, Unknown twice.
- Claude Opus 5.5
- Unknown 6 times.
- Claude Sonnet 5.5
- Unknown 5 times, Nüshu once.
- GPT-6 Astra
- Tibetan twice, Geba once, Marchung once, Unknown once, Marchen once.
- GPT-6.1 Sol
- Tibetan 3 times, Han once, Enochian once, Unknown once.
- Gemini 3.1 Pro
- Unknown 3 times, Paleo-Hebrew twice, Reformed Egyptian once.
- Gemini 3.8 Flash
- Nabataean twice, Linear Elamite once, Proto-Sinaitic once, Unknown once, Tibetan once.
- Grok 4.7
- Arabic twice, No answer once, Unknown once, Georgian Mkhedruli once, Baybayin once.
- Kimi K3
- Unknown 4 times, Latin once, Kharosthi once.
- Llama 4 Maverick
- Mongolian 3 times, Hangul twice, Mkhedruli Georgian once.
- Mistral Medium 3.5
- Rongorongo once, Unknown once, SINHALA once, Sinhala once, Thai once, Syriac once.
- Qwen3.8 Max
- Unknown 5 times, Mongolian script once.
Opus 5.5 checked Tangut, Khitan, Jurchen and Nüshu. It checked Gregg shorthand. It never checked Egypt.
Astra also failed to name the script. We didn't keep the transcript.
They know what a papyrus looks like. They can't read the writing.
Show the best model a real hieratic document, a papyrus, a pottery shard or a plate from an Egyptology book, and it names the script 95% of the time. Show it one sentence in the same script, written fresh by an expert, and it scores 0% across 78 tries.
Ask it which hieroglyph a single hieratic sign stands for and the best model is right 13% of the time. Models seem to know what a page of hieratic looks like. They don't know the signs it is made of.
The goal isn't this sentence. It's hieratic.
The sentence is the exam. Its answer has never been published and isn't stored anywhere, so a model can't pass by remembering something it read online. The only way through is to actually read hieratic.
What we want is AI that can read the handwriting of ancient Egypt, sign by sign, on any papyrus or pottery shard, and help the few people who read it today with the many documents still waiting. A model that can do that will read this sentence too.
Hieroglyphs were for stone. Hieratic was for everything else.
Hieratic is the cursive form of Egyptian hieroglyphs, written fast with a reed brush on papyrus, pottery shards and wooden boards. For more than three thousand years it was how Egypt actually wrote. Letters, tax records, medical manuals, maths problems, love poems, the stories people told.
The Rhind Mathematical Papyrus is hieratic. So is the Edwin Smith surgical papyrus, the Tale of Sinuhe, and the Diary of Merer, a logbook kept during the building of the Great Pyramid and the oldest inscribed papyrus ever found.
Hieroglyphs get the museum walls. Hieratic is where the people are. And almost nobody alive can read it.

Why this is a fair test.
- The answer is sealed.
- The translation has never been published and the benchmark doesn't store it anywhere, so no model can have memorised it.
- Almost nobody can read it.
- Hieratic is read fluently by a small circle of specialists, so there is very little labelled data to learn from. Reading it means learning, not recall.
- Every scribe wrote differently.
- Signs fuse into ligatures and drift across centuries and hands. A model has to generalise from a sign list to a living hand.
- Progress shows up early.
- Naming the script and reading single signs are scored automatically on public data. Partial credit is real credit.
Four rungs, from seeing to reading.
Every rung after the first tells the model it is looking at hieratic, so a model that can't name the script still gets a fair shot at reading it.
1
Identify
What writing system is this?
Asked cold, the way you would ask a friend. Demotic and hieroglyphic controls sit in the set, so answering hieratic every time doesn't pay.
Best so far 95% on real documents, 0% on the sentence
2
Signs
Which hieroglyph is each sign?
Answered in Gardiner sign-list codes, the standard Egyptologists use. Scored by sign error rate.
Best so far 13%
3
Transliterate
How does it sound?
Standard Egyptological transliteration. Not scored on the sentence, since no answer is stored.
Not scored yet
4
Translate
What does it say?
English. Not scored on the sentence either. Scoring is built in for future public texts.
Not scored yet
Leaderboard
Every score comes from the open harness, through Anthropic's API for Claude and OpenRouter for everyone else. The first two columns ask a model to name the script, on the professor's sentence and on real ancient documents. Updated 5 October 2026.
1
Claude Opus 5.5 (high)
- Sentence
- 0%
- Documents
- 95%
- Signs
- 13%
2
Claude Fable 5.1 (high)
- Sentence
- 0%
- Documents
- 95%
- Signs
- 8.7%
3
Claude Sonnet 5.5 (high)
- Sentence
- 0%
- Documents
- 94%
- Signs
- 11%
4
Kimi K3 (high)
- Sentence
- 0%
- Documents
- 84%
- Signs
- Not run
5
Qwen3.8 Max (high)
- Sentence
- 0%
- Documents
- 81%53 of 87
- Signs
- Not run
6
GPT-6.1 Sol (high)
- Sentence
- 0%
- Documents
- 79%
- Signs
- Not run
7
Llama 4 Maverick
- Sentence
- 0%
- Documents
- 61%
- Signs
- Not run
8
Mistral Medium 3.5 (high)
- Sentence
- 0%
- Documents
- 46%
- Signs
- Not run
9
Claude Haiku 4.5
- Sentence
- 0%
- Documents
- 43%
- Signs
- 0.7%
10
Grok 4.7 (high)
- Sentence
- 0%
- Documents
- 25%
- Signs
- Not run
11
GPT-6 Astra (high)
- Sentence
- 0%
- Documents
- Not run
- Signs
- Not run
12
Gemini 3.1 Pro (high)
- Sentence
- 0%
- Documents
- Not run
- Signs
- Not run
13
Gemini 3.8 Flash (high)
- Sentence
- 0%
- Documents
- Not run
- Signs
- Not run
| Rank | Model | The sentence | Real documents | Single signs |
|---|---|---|---|---|
| 1 | Claude Opus 5.5 (high) | 0% | 95% | 13% |
| 2 | Claude Fable 5.1 (high) | 0% | 95% | 8.7% |
| 3 | Claude Sonnet 5.5 (high) | 0% | 94% | 11% |
| 4 | Kimi K3 (high) | 0% | 84% | Not run |
| 5 | Qwen3.8 Max (high) | 0% | 81%53 of 87 | Not run |
| 6 | GPT-6.1 Sol (high) | 0% | 79% | Not run |
| 7 | Llama 4 Maverick | 0% | 61% | Not run |
| 8 | Mistral Medium 3.5 (high) | 0% | 46% | Not run |
| 9 | Claude Haiku 4.5 | 0% | 43% | 0.7% |
| 10 | Grok 4.7 (high) | 0% | 25% | Not run |
| 11 | GPT-6 Astra (high) | 0% | Not run | Not run |
| 12 | Gemini 3.1 Pro (high) | 0% | Not run | Not run |
| 13 | Gemini 3.8 Flash (high) | 0% | Not run | Not run |
Claude Fable 5.1 (high) declined 6 questions, which count as wrong. Real documents counts hieratic documents only, not the demotic and hieroglyphic controls. Transliterating and translating the sentence aren't scored, because no answer is stored anywhere. Chat-app transcripts above are quoted for the record and never counted here.
The dataset
268 items. 87 real hieratic documents, 29 controls in other Egyptian scripts, 150 single signs and 2 sealed sentences. Every public image is openly licensed and credited.

The professor's sentence
2022
Sealed© Aly Moursy. Free to use for evaluating models.

The scribe's rendition
2022
Sealed© Aly Moursy. Free to use for evaluating models.

Papyrus Berlin 3022 (Story of Sinuhe), opening section
Ägyptisches Museum und Papyrussammlung, Berlin
HieraticCC BY-SA 4.0

Ostracon with Pharaoh Spearing a Lion and a Royal Hymn on its Back, Met 26.7.1453
ca. 1200–1080 BCE, The Metropolitan Museum of Art, New York
HieraticCC0 (Met Open Access)

Rhind Mathematical Papyrus (BM EA 10057), detail
c. 1650 BCE, British Museum, London (EA 10057)
HieraticPublic domain

Ostracon: letter from the scribe Mose to the vizier Pesiur (Ashmolean HO 71)
13th century BCE, Ashmolean Museum, Oxford (HO 71)
HieraticCC BY-SA 4.0

Marriage Contract, Met 35.4.1a, b
380–343 BCE, The Metropolitan Museum of Art, New York
DemoticCC0 (Met Open Access)

Single hieratic sign, London, British Museum, EA 9999 (AKU HT 10440)
New Kingdom, Dynasty 20, Ramesses III, regnal year 32, month 3, season šm.w, day 6
Single signCC BY 4.0
Help AI learn to read hieratic.
This only works with more people. It needs the few who can read hieratic, and the people building the models that one day might.
- Egyptologists
- Write us a sentence. Every new sealed sentence makes the benchmark harder to game and the result harder to argue with. Sign annotations on published papyri help just as much.
- Offer a sentence
- AI researchers
- Run the harness on your model with your own key, or build a reader from scratch. Every result goes on the public leaderboard.
- Run the benchmark
- Everyone else
- Share it. Introduce us to an Egyptologist. Sponsor a commissioned sentence. The benchmark grows one sentence at a time.
- Get in touch
Or go see it for yourself.
Egypt's museums are full of hieratic, on papyrus and on pottery shards. Deir el-Medina, near Luxor, is the village where the workmen who built the royal tombs left thousands of notes in it.
Need a visa? Veeza AI handles the Egypt e-visa before you fly.
来源:Hacker News · AI · hieraticbench.vercel.app

