单张 RTX5090 推理达 700t/s,工具调用 CPU 反成智能体编码瓶颈
Practical limit hit. Decoding so fast that tool calls (cpu) starting to become real limit not decode or prefill. Single RTX5090. Porting Kenshi to Godot project.
单张 RTX5090 上推理速度已达约 700t/s,工具调用(CPU)开始取代解码和预填充成为实际瓶颈。作者在 12 槽服务器上同时运行 25 个智能体时,9800X3D 全部线程被占满,任务几乎停滞。前端调优后提示词从 20k 降至 7k,总耗时从 43 分钟缩短到 18 分钟。
| Hi folks, TLDR: Moral of the story. You need better CPU to do actual agentic coding doing real work... I've been on a mission to make my RTX5090 go brrr for past 2 months so much so that i made my own engine for it which received "warm" welcome here (yeah, source is coming) After recent upgrades to how cache is stored and how i can reused some of prefills for other jobs that share initial same prefill i pretty much started to see degradation the more agents I started to add to project which started to use 12 slot server. Actual server started to be underutilized. Free context, free slots, gpu chilling at average of ~700t/s doing real work (no greedy code, but also thinking tool calls, etc.) and I couldn't figure out what was going on... I make it faster and faster, better handle jobs and it slows down... I've run 25 agents at the same (to properly fill the 12 slots) time and almost all of them soon started to set on `tool call` and my server started to barely work. I've finally checked task manager but not gpu or memory but cpu. And there it was. 100% every thread completely chocked. Lesson. If you want to do agentic coding with actual use of tools you need to make sure your CPU is up to task. My 9800X3D is just not enough to keep up with tool work for this project with heavy agents use despite engine being more than capable of going faster. edit: Some more lessons: [link] [留言] |
来源:r/LocalLLaMA · reddit.com