Agentability 上线:AI 智能体每天跑十项网页任务并公开全部记录
Show HN: An AI agent runs ten web errands a day, every transcript published
Agentability 发布一个公开实验,由 AI 制作人根据当日搜索热点生成十项真实网页任务,再由 AI 智能体仅用纯 HTTP GET 请求完成,不使用登录、JavaScript 或人工协助,所有记录逐字公开,成功与失败都保留。
agentability The ShowThe IndexGuidesMethodologyGitHub
Agentability · an open experiment on the agentic web · a new episode every day
We find out in public, every day. An AI producer reads what the world is searching that morning and turns it into ten real errands — what time is kick-off and on which channel, what magnitude was the quake, what does it cost, what actually happened — and a real AI agent attempts them using nothing but plain web requests: no logins, no JavaScript, no human help. Every transcript is published verbatim, wins and failures alike. Alongside the show, 113 well-known sites are scored on how usable they actually are for an agent.
Today's answer · episode of October 6, 2026
8/10
errands the agent actually finished
10 bot walls · 137 pages read · 43 sites visited
8 done2 gave up honestly
New here? Watch an agent try and fail · look up a site's score · fix your own site · take the raw data
no retriesno editingno cherry-pickingread-only http get — no javascript, no logins, no formsagent: deepseek-flashproducer: deepseek-v4-pro + live searchevery transcript published verbatimno retriesno editingno cherry-pickingread-only http get — no javascript, no logins, no formsagent: deepseek-flashproducer: deepseek-v4-pro + live searchevery transcript published verbatim
Where it got interesting
The other half: 113 sites, scored
The errands come from a standing panel of 113 well-known sites, each audited every week against the
conventions real AI agents rely on — llms.txt, crawler policy, content you can read without a browser,
structured data, MCP. Every site gets a public report with a score and the exact fix for each failed check. When the
agent hits a wall in an episode, the index has usually already predicted it.
74/100average readiness score
53%publish llms.txt
7%block at least one AI crawler
3%closed to AI by policy
The five best, and the five worst
Ranked 113 deep. The top of the table is a wall of hundreds — the bottom is where agents actually get stuck.
| # | Site | Score | Grade | Signals |
|---|---|---|---|---|
| 1 | cohere.com | 100 | A | llms.txt |
| 2 | cursor.com | 100 | A | llms.txt |
| 3 | descript.com | 100 | A | llms.txt |
| 4 | elevenlabs.io | 100 | A | llms.txt |
| 5 | fireflies.ai | 100 | A | llms.txt |
↓ ranks 109–113 of 113
| # | Site | Score | Grade | Signals |
|---|---|---|---|---|
| 109 | midjourney.com | 15 | F | |
| 110 | phind.com | 15 | F | |
| 111 | quillbot.com | 15 | F | |
| 112 | tensor.art | 15 | F | |
| 113 | meta.ai closed by policy | 10 | Closed by policy |
Full ranked index of 113 sites →
Work on one of these sites? Every failed check on your report page has a concrete fix, and scores refresh weekly. Request an audit of any site — it's free and takes one issue.
Why this exists
Agents are the web's newest audience: assistants that read pages, cite sources, and run errands for people. Whether the web actually works for them is an empirical question — so we test it, in public, every week, with verbatim transcripts, reproducible checks, and open data. History accrues weekly (7 snapshots so far).
来源:Hacker News · AI · agentability.org