Vals AI 在真实经济任务上评测前沿 AI 模型
Testing AI on Real-World Tasks
Vals AI 对全球领先 AI 模型开展独立评测,覆盖金融、软件以及网络安全、递归自我改进、心理健康等前沿风险任务,所有评测自行执行且多数基准为内部自建。其最新报告显示,Claude Opus 5.5 智能体找到两个室温磁性半导体候选材料,平台同时追踪 16 家实验室的性能前沿。
Independent Evaluation, Unbiased Benchmarks
Testing AI on Real-World Tasks
We benchmark the world's leading AI models on economically valuable tasks such as finance, software, and frontier risk like cybersecurity, recursive self improvement and mental health. We run all of our own evaluations and create many of our benchmarks in-house.
•Oct 02, 2026
Latest Reports
Recent benchmark releases and model evaluations.
Oct 04, 2026
Claude Opus 5.5 agents find two room-temperature magnetic semiconductor candidates
![]()
Industry Leaderboard
Model performance on different sections of the economy.
Benchmark data unavailable
Benchmark data not found
Model Performance Over Time
Tracking how foundation models improve with each release
AccuracyTime
•Oct 05, 2026
Viewing 16 lab performance frontiers.
Vals AI in the Media
Press coverage of our benchmarks and evaluations
来源:Hacker News · AI · vals.ai