Nonobench:49 个 LLM 解数织谜题的开放基准,公开且开源
Nonobench: an open benchmark of 49 LLMs on nonogram puzzles, public and open source [P]
Nonobench 是一个公开开源的基准,用数织(picross)谜题测试 49 个 LLM:模型只拿到一次行列线索,须在无工具、每题仅一次尝试的条件下返回完整网格。
| Nonobench measures how well LLMs solve nonograms (picross). Each model gets the row and column clues once and returns the full grid. No tools, one attempt per puzzle. Method: - Standard mode: 30 puzzles from 5x5 to 15x15 (from the Nonograms dataset by Moyà-Alcover, CC BY 4.0). - Hard mode: ten random 20x20s, each checked to have a single solution. Five can't be solved by line logic alone. Random fills avoid picture puzzles that models can guess. - 130 variants across reasoning effort levels, run through OpenRouter and pinned to each lab's own endpoint where possible. Results: - Solve rates drop from 85% (5x5) to 46% (10x10) to 20% (15x15), each model at its best effort level. - GPT-6 Astra solves all 30 Standard puzzles. On Hard mode, Claude Opus 5.5 solves 8 of 10 and 11 of 15 models solve none. - As one 400-character string, most models lost count before the logic got hard, so Hard mode answers an array of 20 row strings rather than a single string. Limitations: one attempt per puzzle, so single results are noisy (95% intervals shown). Site: https://www.nonobench.com Code (MIT): https://github.com/mauricekleine/nonobench submitted by /u/mauricekleine[link] [留言] |
来源:r/MachineLearning · reddit.com