跳到正文
r/LocalLLaMA· /u/Hyungsun·· 3 小时前AI 评分34

M3 Max 36 GB 本地推理实测:oMLX vs Rapid-MLX vs Splash vs MTPLX,Qwen3.6-35B-A3B 达 110 tok/s

oMLX vs Rapid-MLX vs Splash vs MTPLX on M3 Max 36 GB: 110 tok/s on Qwen3.6-35B-A3B, ~32 tok/s on Qwen3.8-27B

AI 导读

在 36 GB M3 Max 上实测 oMLX 0.7.0、Rapid-MLX 0.15.5、Splash 1.2.1、MTPLX 2.12.2 四款本地推理引擎。

正文

Hello. I picked up a new old stock 14" M3 Max MacBook Pro (14 core CPU / 30 core GPU / 36 GB / 1 TB) from my local market yesterday for around $2,498, and spent the night testing which local inference software is actually fastest on it for the two models I use.

Four engines, all current versions: oMLX 0.7.0, Rapid-MLX 0.15.5, Splash 1.2.1, MTPLX 2.12.2. macOS 27 Golden Gate.

Models: Qwen3.8-27B-4bit (dense) and Qwen3.6-35B-A3B-4bit (MoE).

One thing up front: it is not the exact same weight file on all four engines. Rapid and MTPLX run their own MTP-augmented 4-bit builds, Splash pairs its own DFlash2 draft, and oMLX ran the plain mlx-community 4-bit. Same base models, different finishing, but that is how each app is meant to be used.

How I tested: each engine served on localhost, temperature 0, thinking off. Sustained test: same short prose prompt, 3 runs x 256 output tokens, median. Then a prompt size sweep at about 130 / 1500 / 5500 tokens. For thermals I used a laptop stand, waited 3 minutes between every engine+model combo and 2 minutes between the two test phases, and cleared each engine's KV cache before its turn (oMLX, MTPLX and Rapid-MLX all keep caches across restarts, great for daily use, but it will fool you if you benchmark twice). I re-ran the whole thing end to end and the numbers came back within 6%.

Decode, natural prose prompt, median of 3 runs:

engine Qwen3.8-27B Qwen3.6-35B-A3B
Rapid-MLX 27.7 tok/s 110.2 tok/s
oMLX 17.9 tok/s 104.2 tok/s
Splash 31.7 tok/s 77.0 tok/s
MTPLX 31.5 tok/s 79.0 tok/s

Same thing with filler prompts at longer sizes (repetitive text makes speculative decoding look better, so read this as a best case):

engine 27B @ 1.5K MoE @ 1.5K MoE @ 5.5K
Rapid-MLX 31.3 tok/s 118.2 tok/s 117.9 tok/s
oMLX 17.9 tok/s 101.1 tok/s 95.6 tok/s
Splash 52.2 tok/s 238.0 tok/s 100.9 tok/s
MTPLX 30.2 tok/s 78.6 tok/s 69.8 tok/s

What I take from it:

  • Absolute fastest per model: the MoE goes to Rapid-MLX, dense goes to Splash (31.7 vs MTPLX 31.5, in practice a tie). oMLX is way behind on dense at 17.9 but basically level with the leaders on the MoE.
  • If the margins are too small to care about, just pick by features. Splash and MTPLX are the same on dense, Rapid and oMLX are the same on the MoE. I kept Rapid-MLX because the MoE is my daily model and it is fastest there.
  • The dense number makes sense: roughly 16 GB of weights per token against ~300 GB/s of memory bandwidth puts the ceiling near 18 tok/s, and oMLX sits right on it. The others pass it with speculative decoding, which is also why their numbers move with the kind of text generated. Splash on the MoE was 238 tok/s on the filler prompt at 1.5K and 101 tok/s at 5.5K, while Rapid stayed around 118 tok/s.
  • First token on a 5.5K prompt: about 4-5 s on the MoE, ~34 s on the dense (prefill around 1.2-1.3k tok/s vs ~150-170 tok/s).
  • It is loud under sustained inference. Fans stay up while it generates. Works on my desk, would not use it in a library.
  • I also tried Qwen3.8-Flash-Next (the 125B). Not happening on 36 GB. The 4-bit weights alone are ~74-83 GB and the lightest build asks for 96 GB+, and none of these engines can stream that architecture's experts off the SSD.

Limitations: one laptop, one night, medians of 3 runs, and the different weight builds mentioned above. My prompts are simple too, no long agent sessions yet.

TL;DR: on a 36 GB M3 Max, Qwen3.6-35B-A3B does ~110 tok/s on Rapid-MLX and Qwen3.8-27B ~31.7 tok/s on Splash (MTPLX a hair behind), pick by which model you run most, and the 125B Flash-Next needs 96 GB+.

submitted by /u/Hyungsun
[link] [留言]

来源:r/LocalLLaMA · reddit.com