跳到正文
r/LocalLLaMA· /u/swiebertjee·· 7 小时前AI 评分22

双 DGX Spark 用户实测:GLM 5.3 flash 性能提升 50% 以上

For dual DGX spark users; GLM 5.3 flash got a 50%+ performance boost

AI 导读

双 DGX Spark 用户实测显示,GLM 5.3 flash 在新版 recipe 下解码性能提升 50%-90%,解码速度已超过 DeepSeek v4.0 flash,仅 prefill 略有下降。

正文

For the last few months, I ran DeepSeek v4.0 flash (NVFP4). First 0731, then visionexp because it was a free improvement. I got around 65 tps decode and almost 2k prefill, and ran 4-5 agents in parallel, totalling around 200 tps cumulative decode. Because of this, I did not feel like switching to GLM 5.3 because it would half the decode and prefill, did not scale well with multiple agents, and had a repetition bug a lot of people complained about.

Until a few days ago, when the latest version of this recipe dropped; a 50-90% decode improvement. So I took the plunge, and wow, am I impressed.

It's more intelligent than the new DeepSeek v4.1 flash (that does NOT run on dual DGX Sparks), and it's even faster than DeepSeek v4.0 flash in decode. Only a slight drop in prefill, which I'm more than happy to take in exchange;

Test visionexp-final (recorded) glm53-low Δ
B1 count-to-300 92.5 95.9 +4%
B1 bulk SQL INSERT 88.3 97.2 +10%
B2 chat 38.7 42.5 +10%
B2 count 92.8 96.0 +3%
B2 code 63.5 69.3 +9%
B2 prose 32.8 37.1 +13%
B2 tool 79.8 85.1 +7%
B2 battery mean 61.5 66.0 +7%
B2 accepted tok/step 3.26 of 6 (54%) 3.55 of 8 (44%) see note
B3 prefill @1.5K 1738 1376 -21%
B4 prefill @32K 1902 1576 -17%
B4 prefill @128K 1758 1578 -10%
B4 decode @32K 41.1 44.2 +7%
B4 decode @128K 49.5 47.1 -5%
B5 c1 aggregate 91.7 90.1 -2%
B5 c2 aggregate 45.4 51.4 +13%
B5 c4 aggregate 63.4 58.6 -7%
B5 c6 aggregate 79.1 77.5 -2%
B7 soak (40 min at c4) 522 req, 0 err, 87.4 agg 503 req, 0 err, 0 soft-empty, 83.6 agg −4%
B8 byte-stable probes 8/8 6/8 worse
B8 garble gate 30/30 clean 30/30 clean =
B8 non-Latin / U+FFFD not measured 3/3 clean, 0 U+FFFD new gate
KV pool 1,988,929 tok @ gmu 0.85 560,362 tok (6 GiB/rank pin) −72%
NRestarts through the pass 0 0 =

I've tested it for a few days now, both for technical coding, devops/sysadmin and also vision (to recognize some plants), and it is better than I hoped for. Basically Claude Opus 4.8 level. Slower of course because it has to think a lot more, but good enough to comfortably leave it chugging for hours on tickets without worry of derailing. I don't see a reason NOT to upgrade, so have a try and enjoy!

submitted by /u/swiebertjee
[link] [留言]

来源:r/LocalLLaMA · reddit.com