跳到正文
r/LocalLLaMA· /u/pand5461·· 20 小时前AI 评分22

同一 IQ3_S 模型在 llama.cpp 中不再陷入死循环

Need maybe say "Use llama.cpp"

AI 导读

有用户测试 IQ3_S 模型时发现,该模型在长文本生成中约 1 万 token 后陷入重复输出的死循环,反复念叨"Need maybe"之类的句子。同一模型在 llama.cpp 中却能给出连贯回答,未出现该问题。

正文

So I tried that miracle engine everyone is talking about.

Asked the IQ3_S model to express its opinion on a post from this sub to measure the tps on a long-ish generation:

Can you help with the following problem?

So Kimi K2 is outdated, and so is GPT OSS 120b. Which of the modern open weights models can boast the least sycophancy? I need this both for creative/research assistant usage (sycophancy led me down blind alleys of my own bad ideas many times) and agentic coding (more sycophancy less bug noticing).

The thinking trace:

We need answer user's question. Need likely provide current landscape as of 2026? We have get_datetime tool. Need know current date 2026? System says current date 2026-06-22. Need maybe use get_datetime? Could call to confirm. User asks about modern open weights models least sycophancy. Need likely discuss Kimi K2 outdated? ... 

10k tokens later it degrades to:

Need maybe maybe include "Use 'for code, list constraints'." Need maybe maybe include "Use 'for code, list requirements'." 

The same exact model in llama.cpp does produce a coherent answer without a doom loop.

submitted by /u/pand5461
[link] [留言]

来源:r/LocalLLaMA · reddit.com