跳到正文
elvis· @omarsar0 · X·· 4 小时前AI 评分56
AI 导读

Tinker 宣布提升效率以支持长上下文 RL 扩展,价格下调最高 70%,长上下文 prefill 与采样现在与短上下文同价,GLM-5.3-Flash 和 DeepSeek-v4.1-Flash 也已上线用于低成本长上下文任务。

正文

Bullish on this trend of making post-training more accessible.

A new post-training era is upon us.

If you work on agentic RL, long-context tasks (a big focus today) are expensive, inefficient, and don't scale well.

I've been diving into RL envs and evals for long-context tasks, and I can see this being useful.

In agent RL, rollouts use most of the tokens. Every turn re-reads the whole growing context, including tool outputs, files, and earlier turns.

Tinker just cut the price of those tokens. Long-context prefill and sampling now cost the same as short context.

This means that evaluating your trained models on long inputs also gets cheaper. Huge win here.

I believe RL will keep unlocking specialized models that slash the cost of critical agent operations. Cheaper long rollouts make them more practical to build.

Own your intelligence stack!

引用Tinker@tinkerapi
Tinkerers have been busy scaling up long-context RL! We’ve made significant improvements to Tinker’s efficiency to support those, and are passing these on with price cuts up to 70%. GLM-5.3-Flash and DeepSeek-v4.1-Flash are also live for cost-efficient long-context work.
在 X 查看被引用的帖子

来源:elvis · x.com