跳到正文
r/LocalLLaMA· /u/khiladi796·· 7 小时前AI 评分34

小型推理模型(SRM)会是下一个大转向吗?我们到底该测什么?

Are "small reasoning models" the next big shift? What should we actually be measuring?

AI 导读

有开发者提出,小型推理模型(SRM)可能凭借原生潜空间推理,在不吸收海量通用知识的情况下逼近成本-准确率帕累托前沿,Pathway 的 ARC-AGI-1 结果被视为一个例证。讨论建议从紧凑度(实际推理内存与算力)、few-shot 适配方式、训练数据效率、持续学习四项指标评估这类模型,而非只看营销说法。BDH 等架构因状态与记忆处理方式不同于标准 Transformer,被指可避免灾难性遗忘。

正文

For a model running locally on a fairly narrow task, how much general knowledge do we actually need, and how much reasoning capability could we get without it ?

SRMs are interesting for obvious reasons, but I went down this rabbit hole after listening to Ben Lorica's (advisor at Databricks) chat with Zuzanna Stamirowska from Pathway (BDH). Ben keeps coming back to this broader theme of how "specialized AI is getting easier to build" and the Kumo RFM angle, but it opened an interesting thread around small reasoning models.

Pathway’s ARC-AGI-1 result makes an interesting case for small models hitting the cost-accuracy Pareto frontier. The premise is that if a model is built to reason natively in its latent space, it might not need billions of parameters absorbing Reddit and Wikipedia just to solve logic puzzles. They described a use-case of long-horizon reasoning within a bounded domain as a target (like security investigations, tickets analysis – a real case I know from a major bank, etc.)

It's obvious that just because a large model does well in 20 languages. I don't need that for work tasks. There is definitely a market for compact models with substantial reasoning ability.

Also because architectures like BDH handle state and memory differently than standard transformers, the pitch is that they avoid catastrophic forgetting (learning continuously from new examples at inference time without wiping past skills). The question is how to evaluate this without getting lost in marketing claims. Here is how I'd break it down

• Compactness: Low parameter count, but what are the actual inference memory and compute requirements?

• Few-shot adaptation: Does it adapt through context or actual parameter updates?

• Training efficiency: How much data did it actually need to pick up the underlying capability?

• Continual learning: Does post-deployment experience produce persistent improvements without degrading earlier skills?

For people running small models locally: what workload expose the difference between a compact model that just follows in-context examples versus one that actually learns reusable rules?

submitted by /u/khiladi796
[link] [留言]

来源:r/LocalLLaMA · reddit.com