DeepSeek API News·· 2025-05-28精选AI 评分64
DeepSeek 将 deepseek-reasoner 升级至 DeepSeek-R1-0528
deepseek-reasoner
AI 导读
DeepSeek 将 deepseek-reasoner 模型升级至 DeepSeek-R1-0528,多项基准 Pass@1 提升:AIME 2025 从 70.0 升至 87.5,GPQA 从 71.5 升至 81.0,LCB_v6 从 63.5 升至 73.3,Aider 从 57.0 升至 71.6。
推荐理由
DeepSeek-R1-0528 在 AIME、GPQA、Aider 等基准上的提升幅度,可用于判断推理模型迭代节奏。
正文 · 原文
deepseek-reasoner Model Upgraded to DeepSeek-R1-0528:
- Enhanced Reasoning Capabilities
- Significant benchmark improvements (Pass@1)
- AIME 2025: 70.0 → 87.5 (+17.5)
- GPQA: 71.5 → 81.0 (+9.5)
- LCB_v6: 63.5 → 73.3 (+9.8)
- Aider: 57.0 → 71.6 (+14.6)
- Note: Complex reasoning tasks may consume more tokens compared to legacy R1 version.
- Significant benchmark improvements (Pass@1)
- Optimized Front-end Development
- Generated web pages and games now feature improved aesthetics.
- Reduced Hallucinations
- Significantly suppressed hallucination issues present in legacy R1 version.
- JSON Output & Function Calling Support
- Function call performance:
- Tau-bench score: 53.5 (Airline) / 63.9 (Retail)
- Function call performance:
来源:DeepSeek API News · api-docs.deepseek.com