跳到正文
Hacker News · AI· khalidsaidi·· 4 小时前AI 评分56

Agentability 上线:AI 智能体每天跑十项网页任务并公开全部记录

Show HN: An AI agent runs ten web errands a day, every transcript published

AI 导读

Agentability 发布一个公开实验,由 AI 制作人根据当日搜索热点生成十项真实网页任务,再由 AI 智能体仅用纯 HTTP GET 请求完成,不使用登录、JavaScript 或人工协助,所有记录逐字公开,成功与失败都保留。

正文

agentability The ShowThe IndexGuidesMethodologyGitHub

Agentability · an open experiment on the agentic web · a new episode every day

We find out in public, every day. An AI producer reads what the world is searching that morning and turns it into ten real errands — what time is kick-off and on which channel, what magnitude was the quake, what does it cost, what actually happened — and a real AI agent attempts them using nothing but plain web requests: no logins, no JavaScript, no human help. Every transcript is published verbatim, wins and failures alike. Alongside the show, 113 well-known sites are scored on how usable they actually are for an agent.

Today's answer · episode of October 6, 2026

8/10

errands the agent actually finished

10 bot walls · 137 pages read · 43 sites visited

8 done2 gave up honestly

New here? Watch an agent try and fail · look up a site's score · fix your own site · take the raw data

no retriesno editingno cherry-pickingread-only http get — no javascript, no logins, no formsagent: deepseek-flashproducer: deepseek-v4-pro + live searchevery transcript published verbatimno retriesno editingno cherry-pickingread-only http get — no javascript, no logins, no formsagent: deepseek-flashproducer: deepseek-v4-pro + live searchevery transcript published verbatim

Where it got interesting

The other half: 113 sites, scored

The errands come from a standing panel of 113 well-known sites, each audited every week against the conventions real AI agents rely on — llms.txt, crawler policy, content you can read without a browser, structured data, MCP. Every site gets a public report with a score and the exact fix for each failed check. When the agent hits a wall in an episode, the index has usually already predicted it.

74/100average readiness score

53%publish llms.txt

7%block at least one AI crawler

3%closed to AI by policy

The five best, and the five worst

Ranked 113 deep. The top of the table is a wall of hundreds — the bottom is where agents actually get stuck.

#SiteScoreGradeSignals
1 cohere.com 100 A llms.txt
2 cursor.com 100 A llms.txt
3 descript.com 100 A llms.txt
4 elevenlabs.io 100 A llms.txt
5 fireflies.ai 100 A llms.txt

↓ ranks 109–113 of 113

#SiteScoreGradeSignals
109 midjourney.com 15 F
110 phind.com 15 F
111 quillbot.com 15 F
112 tensor.art 15 F
113 meta.ai closed by policy 10 Closed by policy

Full ranked index of 113 sites →

Work on one of these sites? Every failed check on your report page has a concrete fix, and scores refresh weekly. Request an audit of any site — it's free and takes one issue.

Why this exists

Agents are the web's newest audience: assistants that read pages, cite sources, and run errands for people. Whether the web actually works for them is an empirical question — so we test it, in public, every week, with verbatim transcripts, reproducible checks, and open data. History accrues weekly (7 snapshots so far).

来源:Hacker News · AI · agentability.org