Tinker 宣布提升效率以支持长上下文 RL 扩展,价格下调最高 70%,长上下文 prefill 与采样现在与短上下文同价,GLM-5.3-Flash 和 DeepSeek-v4.1-Flash 也已上线用于低成本长上下文任务。
Bullish on this trend of making post-training more accessible.
A new post-training era is upon us.
If you work on agentic RL, long-context tasks (a big focus today) are expensive, inefficient, and don't scale well.
I've been diving into RL envs and evals for long-context tasks, and I can see this being useful.
In agent RL, rollouts use most of the tokens. Every turn re-reads the whole growing context, including tool outputs, files, and earlier turns.
Tinker just cut the price of those tokens. Long-context prefill and sampling now cost the same as short context.
This means that evaluating your trained models on long inputs also gets cheaper. Huge win here.
I believe RL will keep unlocking specialized models that slash the cost of critical agent operations. Cheaper long rollouts make them more practical to build.
Own your intelligence stack!
Tinkerers have been busy scaling up long-context RL! We’ve made significant improvements to Tinker’s efficiency to support those, and are passing these on with price cuts up to 70%. GLM-5.3-Flash and DeepSeek-v4.1-Flash are also live for cost-efficient long-context work.在 X 查看被引用的帖子
来源:elvis · x.com