Swift 1.5 在 veloGB10 上跑 2× DGX Spark:xhigh 档约 110 tok/s,速度反超 vLLM 上的 base Flash-Next medium
Swift 1.5 on veloGB10, ~110 tok/s on 2× DGX Spark: xhigh beats base Flash-Next at medium on vLLM
UkisAI 对 Qwen3.8-Flash-Next 的推理精简微调 Swift 1.5,在 sf-stav 专为 GB10 打造的 Rust/CUDA 引擎 veloGB10 上以 xhigh 档运行,单流解码约 110 tok/s,日常编码任务每轮 371 秒,快于旧 vLLM NVFP4 配置下 base Flash-Next medium 的 506 秒。
On my two DGX Sparks, Swift 1.5 (UkisAI's reasoning-efficient fine-tune of Qwen3.8-Flash-Next) running on veloGB10 (https://github.com/sf-stav/veloGB10, sf-stav's Rust/CUDA engine built only for GB10) lets me run coding agents at xhigh effort and still finish sooner than base Flash-Next at medium did on my old vLLM NVFP4 setup.
| Metric | Base Flash-Next NVFP4 @ medium, vLLM | Swift 1.5 EXL3 @ xhigh, veloGB10 |
|---|---|---|
| Single-stream decode | ~52 tok/s | ~110 tok/s |
| Everyday coding tasks, time per pass (5 tasks) | 506 s | 371 s |
| Hard trap tasks, time per pass (4 tasks) | 714 s | 689 s |
| Hard trap tasks, pass rate | 50% (1 pass × 4 tasks) | 92% (3 passes × 4 tasks) |
Both columns run the same agentic battery: Claude Code driving the model through real tool use in a copy of a real repo, graded by test oracles (details below). One note on that score row: the same base model at medium scored 10/12 on velo (table further down), so most of the vLLM score gap is that older setup, not the model (I am rerunning this right now for an even comparison on intelligence but would expect it to be quite similar to the below medium results).
The speed is the point: on velo, xhigh fits in the time medium used to take. However, velo can't load Swift, or any other community EXL3 pack of Flash-Next I could find, out of the box. The fix is a header-only rewrite below.
Caveats: The vLLM numbers are from September: a single pass, on an older version of my serving setup, not a same-day rerun (I've since moved the worker to velo). Most of the speed is velo's: about 2× the decode rate is what pays for xhigh's extra thinking. How much of the score comes from Swift and how much from xhigh itself I can't separate yet, but a base-weights run at xhigh is going now and I'll add it as an update. Twelve runs is still a small sample regardless.
I looked first: everything published about Velo uses one pack, the official doth4580 EXL3, and I couldn't find anyone here, on the NVIDIA forums or in the repo's issues, running Swift, or any other fine-tune, on it.
What breaks
The first community pack I tried (alesha-pro/Huihui-Qwen3.8-Flash-Next-abliterated-exl3-4bit-hq_h6_ng6) died at boot with ple shard 0 not in index. Swift 1.5's EXL3 builds ship the same layout: current exllamav3 (1.5.x) writes the model's 51B-parameter n-gram table as 128 shard tensors in ngram_embedding.safetensors, which isn't listed in the index. velo reads either one big tensor (the doth4580 layout) or indexed 5-bit (K5) shards only, and the 4.05 packs use 6-bit (K6).
Which packs this affects
I read the n-gram header of every Flash-Next EXL3 pack I could find (HTTP range requests on the headers, nothing downloaded):
| Pack | n-gram layout | velo v0.7.2 |
|---|---|---|
| doth4580 4.05 / turboderp 4.05 | one tensor | loads as shipped |
| Swift 1.5: SharkWipf 4.05 | 128 contiguous shards | needs the fix — tested, works |
| Swift 1.5: KatterMobile 4.05, SharkWipf 5.52, scorpoon 3.25, thelastspark 4.00 / 6.05 | 128 contiguous shards | needs the fix |
| Huihui abliterated (alesha-pro 4.05) | 128 contiguous shards | needs the fix — tested, works |
| heretic 3.05 (andrevp, jeffpeng3), Uncensored 4.0 (Lygodactylus), groxaxo 3.50, turboderp 3.05 | 128 contiguous shards | needs the fix |
12 of the 14 builds need it, including all six Swift 1.5 builds.
The fix
In every pack I checked, the 128 shards sit back to back in order. So rewriting only the safetensors header to describe them as one tensor over the same bytes makes velo's single-tensor path load them. No data is copied, the header stays the same length, and the original header is saved for rollback. Script and details: https://github.com/sf-stav/veloGB10/issues/9
Results (TP=2, both Sparks, after the fix)
| Metric | Official doth4580 4.05 | Huihui abliterated 4.05 | Swift 1.5 (SharkWipf 4.05) |
|---|---|---|---|
| Single-stream decode | 110.6 tok/s | 113.6 tok/s | 110.2 tok/s |
| Sanity set (chat, code, JSON, tool call, 38.9K recall) | 5/5 | 5/5 | 5/5 |
| MTP draft acceptance | 62–86% | 44–84% | 33–77% |
All three run at the same speed. The fix is only a header change, so nothing about the weights or kernels differs.
Does Swift actually think less on velo? On the 22 test prompts that ship with the doth4580 pack, run 3 times each (rendered at medium effort, same sampler on both), Swift 1.5 generated 8.2% fewer tokens than the official model: fewer on 16 of 22 prompts, median −8.6% per prompt. Thinking's share of the output fell from 48% to 42%, and total time fell 11%. That's real but far below UkisAI's 63% headline, which was measured at high effort, where there's much more overthinking to remove. At medium, the base model already keeps its thinking short.
Does it still code? I run a private agentic coding battery: Claude Code driving the model through real tool use in a copy of a real repo, graded by test oracles. Each was run 3 times:
| Metric | Official 4.05 | Huihui abliterated | Swift 1.5 | Swift 1.5 @ xhigh |
|---|---|---|---|---|
| Effort | medium | medium | medium | xhigh |
| Everyday tasks (5 tasks × 3) | 15/15 | 15/15 | 15/15 | 15/15 |
| Hard trap tasks (4 tasks × 3) | 10/12 | 7/12 | 7/12 | 11/12 |
| Wall time per everyday pass | —* | 232 s | 199 s | 371 s |
| Wall time per hard pass | —* | 381 s | 375 s | 689 s |
| Output tokens, everyday ×3 | —* | 56K | 49K | 100K |
*The official pack's runs hit a streaming bug in my proxy setup that roughly doubled their wall time, so I've left its times and tokens out. Its pass/fail results are unaffected.
At medium, everyday coding is identical across all three. On the hard set (tasks built from real failures: a brief that states something false, a code review with one planted wrong finding, and so on) both fine-tunes score 7/12 against 10/12 for the official pack, with their misses on the same tasks. I checked that the model received byte-for-byte the same request parameters in both runs, so it's not the harness.
Swift at xhigh went from 7/12 to 11/12, the best result I've had on this battery from any model, at about 2× the output tokens and 1.8× the wall time of medium. The failures it stopped making are the expensive ones in practice: leaving a sibling test suite broken without saying so, acting on the planted wrong review finding and going out of scope to do it, and an off-by-one in a date window. The failure was the "mirror" trap: asked to add a new league by following an existing one, it copied tuning values the new league doesn't have data for.
That's the hardest task demonstrated, and Swift at xhigh passed it 2 times out of 3 while no other configuration in the table passed it more than once. On velo, all that extra thinking still lands inside the time base Flash-Next at medium took on vLLM (the table at the top).
12 runs per configuration is a small sample, as mentioned before (Fisher p ≈ 0.4 for 10 vs 7, ≈ 0.15 for Swift xhigh vs Swift medium), so "suggestive," not proven. I haven't run the official or abliterated packs at xhigh yet, so I can't yet tell how much of that jump is Swift and how much is just the higher effort. velo's loop detector was off for all battery runs.
Baseline numbers (official pack, TP=2)
- Single-stream decode: 110.6 tok/s (vs ~52 on my vLLM NVFP4 setup, measured in September, not same-day).
- Time to first token at 4K / 16K / 64K: 2.0 / 7.4 / 25.7 s.
- Concurrency is the catch: at 2–3 requests they take turns (aggregate 91 → 95 → 98 tok/s); from ~4 they batch (157 at 8, 181 at 16) but I saw the author say they were working on it this week.
Gotchas
- TP=2 with the cable on the f0 ports:
--rdma-dev rocep1s0f0,roceP2p1s0f0(velo defaults to f1). - llama-benchy's prefill t/s is wrong for velo (first SSE event arrives before prefill); use e2e TTFT.
- OpenAI chat/completions only, no
/v1/responses: use litellmhosted_vllm/, notopenai/. --model-nameis ignored on the EXL3 path; the model id is the pack's folder name.
Credit: sf-stav (veloGB10), turboderp (exllamav3), doth4580, UkisAI (Swift), huihui-ai and every quant uploader in the table. I've filed the loader issue upstream (https://github.com/sf-stav/veloGB10/issues/9) so packs can eventually load as shipped.
submitted by /u/MushroomMan234
[link] [留言]
来源:r/LocalLLaMA · reddit.com