跳到正文
r/LocalLLaMA· /u/SeriousJul·· 5 小时前AI 评分36

Qwen3.8 27B 对比 Flash Next:基准测试之外的实际体验

Qwen3.8: 27b vs flash next. We all know the benchmarks, but at least to me, the reality is a different story

AI 导读

用户实测 Qwen3.8-27B(Q4_K_XL 量化、130K 上下文)与云端 Flash Next(256K 上下文)对比,尽管基准测试中 Flash Next 略优,但 27B 在简单工作流中质量更好,约 2 次迭代即可完成,而 Flash Next 需约 5 次。27B 的 PR 评论更简洁直接,最终代码质量相当,但 Flash Next 存在过度冗长和深度幻觉问题。

正文

By classic benchmark, the flash next is supposed to be slightly superior to its dense counterpart. But they are really incomparable. For my very simple workflows (spec -> implement -> review <-> rework), I feel that 27b is just better quality.

For context, and making things worse, I am comparing quantized 27b versus cloud flash next.

- self hosted unsloth/Qwen3.8-27B-GGUF:Q4_K_XL (stock llamacpp with 130K context window)
- alibaba cloud (qwen individual token plan), context capped at 256K in the harness

The metrics for my quality is actually very simple, I measure the number of review / rework needed before a PR is ready for me to read. The tasks are all very simple with a tight scope. Usually 27b do the work in ~2 iterations, flash next needs ~5. And it is not only about the number of iteration.
On the review flash next is overly verbose on half baked PR comment, where 27b is more straight to the point. In the end, the code produced is on par, to be frank. But if we look at token consumption...

Side notes, on pairing session, I got some deep hallucination using "/skills:diagnosing-bugs" + flash next. But since it is "bugs" and they not really comparable chunk of work, it is hard to say.

And I can't be the only one feeling that right ? Are you feeling the same ?

PS: of course I followed the hype and jumped on Strata. After the initial "oh my god it's so fast", I switched back to 27b. Tried all quant from "ISTA-DASLab" as well as experiemental from unsloth (Q4_K_L). With ISTA-DASLab, It actually is the first time I had "tool call error" in pi (which stop the agent), multiple times.

submitted by /u/SeriousJul
[link] [留言]

来源:r/LocalLLaMA · reddit.com