跳到正文
Hacker News · AI· doener·· 4 小时前AI 评分22

Vals AI 在真实经济任务上评测前沿 AI 模型

Testing AI on Real-World Tasks

AI 导读

Vals AI 对全球领先 AI 模型开展独立评测,覆盖金融、软件以及网络安全、递归自我改进、心理健康等前沿风险任务,所有评测自行执行且多数基准为内部自建。其最新报告显示,Claude Opus 5.5 智能体找到两个室温磁性半导体候选材料,平台同时追踪 16 家实验室的性能前沿。

正文

Independent Evaluation, Unbiased Benchmarks

Testing AI on Real-World Tasks

We benchmark the world's leading AI models on economically valuable tasks such as finance, software, and frontier risk like cybersecurity, recursive self improvement and mental health. We run all of our own evaluations and create many of our benchmarks in-house.

•Oct 02, 2026

Latest Reports

Recent benchmark releases and model evaluations.

Oct 04, 2026

Claude Opus 5.5 agents find two room-temperature magnetic semiconductor candidates

Two bar magnets pointing opposite ways, one green and one gray, with their field lines looping into each other

Industry Leaderboard

Model performance on different sections of the economy.

Benchmark data unavailable

Benchmark data not found

Model Performance Over Time

Tracking how foundation models improve with each release

AccuracyTime

Vals Index•Oct 05, 2026

Viewing 16 lab performance frontiers.

Vals AI in the Media

Press coverage of our benchmarks and evaluations

View All News

来源:Hacker News · AI · vals.ai