清华大学一篇新论文提出 TokenRouter,一个在 token 级别对大小模型进行路由的服务系统,吞吐量最高达到现有方案的 64.15 倍。现有主流服务框架如 vLLM 和 SGLang 每个请求只跑一个模型,当两个模型共同生成答案时,每一步都要等较慢的那个。
New Tsinghua paper builds TokenRouter, a serving system that runs per-token small-and-large model routing at upto 64.15X the throughput of existing setups.
Current popular serving frameworks (like vLLM and SGLang) run one model per request, so when two models share an answer, every step waits for the slower one.
TokenRouter gives each model its own server and lets them pass work back and forth. So, it hands a half-written answer between them while keeping the model’s memory of the text so far (the KV cache), and holds requests for a moment so each model works on bigger batches.
Across 5 routing methods, throughput rose 2.01 to 64.15 times over the stronger existing setup.
– arxiv. org/abs/2610.12242
Title: "TokenRouter: Efficient Serving System for Token-Level LLM Routing"
来源:Rohan Paul · x.com