Qwen3.8-Flash-Next 与 Claude Opus 5.5 在 Strix Halo 笔记本上的同一功能实测对比
Story time: Qwen3.8-Flash-Next on my Strix Halo laptop vs Claude Opus 5.5 on the same feature
作者在 ASUS ROG Flow Z13(Strix Halo,128GB)上用 Qwen3.8-Flash-Next(xhigh 档,Pi 作为 harness)与 Claude Code 中的 Opus 5.5(medium 档)为 LlamaStash 这个大型 Rust 项目实现同一个 daemon 重启命令。
For the last few weeks most of my coding has been done locally with Qwen3.8-Flash-Next, so I gave it and Opus 5.5 the same high complexity feature to build on LlamaStash (a complex and large Rust project) and compared the results.
Setup: ASUS ROG Flow Z13 (Strix Halo, 128GB), Flash-Next at xhigh effort with Pi as the harness. Opus 5.5 ran in Claude Code at medium effort. I wanted xhigh for Opus as well, but Claude changed it to medium when I picked the latest model and I didn't notice it until the task was done. But I think medium is probabbly a fairer setting anyway.
Task: add a llamastash daemon restart command that reuses the existing start and stop code. I kept the prompts vague on purpose and gave both the same prompts.
| Step | Opus 5.5 (medium) PR#88 | Flash-Next (xhigh) PR#89 |
|---|---|---|
| First iteration | ~9 min | ~38 min |
| Nudge to reuse the TUI restart code | ~6 min | ~34 min |
| A third duplicate path | found it on its own | ~30 min, after one more prompt |
| Create PR | ~3 min | ~30 min |
| Total | ~18 min | ~130 min |
| Tokens (in / out) | 7.83M / 41.5K | 20.61M / 101K |
| Tests added | 1 | 4 (2 of them end to end) |
| Cost | $7.53 | $0 + ~0.15 kWh |
The end result was interesting. I asked GPT 5.6, Opus 5.5 and Flash-Next to review and compare both PRs (new sessions). GPT and Flash-Next picked the Flash-Next PR (#89) and Opus picked its own (#88). I also did my own review and found the Flash-Next one better as it had better tests and handled edge cases better. I ended up merging #89, after porting the fixes that the reviews picked from #88.
Keep in mind:
- Opus was on medium effort. With xhigh it would have used way more tokens, taken a bit more time and probably would have done a better implementation.
- Flash-Next ran on an older Halogen version (0.14.0), and Halogen dropped the connection once, so the last part ran on Gufo. The current Halogen does around 1,400 t/s prefill and 46 t/s decode on my laptop at 70 W, so I think the time will drop a lot if I redo the test.
- The $7.53 is what Claude Code reported for the whole Opus session, which includes a later fix to the PR. The 0.15 kWh assumes 70 W for the whole 130 minutes.
Opus is still 2 to 10 times faster and I still use it for planning and reviews. But the actual coding now happens on my laptop, and to me it is crazy that I can run a local model that can challenge a frontier model like this.
Full post with my setup, the engine benchmarks and a second task comparison: https://deepu.tech/local-ai-qwen3.8-flash-next-best-local-llm
submitted by /u/deepu105
[link] [留言]
来源:r/LocalLLaMA · reddit.com