MLX-Serve 26.10.1 (main 599dd05, the release candidate) against the shipped MLX-Serve 26.9.6, same session, arms alternating, llmprobe measuring; TensorFold beside them where the same pack was measured on the same box. Decode and prefill in tokens per second. Snapshot taken 2026-10-01 11:13 EDT while the runs continue on four Macs.
MLX-Serve 26.9.6MLX-Serve 26.10.1TensorFold (0.6.0 on every M5 Ultra row, 0.5.0 elsewhere; same box, same llmprobe method; the 6/8-bit rows run its drafter)
Agent build check
Three three.js games by pi on Flash Next (MLX-Serve 26.10.1, M5 Ultra)
one prompt each, single-file index.html, loaded in headless Chromium
Breakout · one pi turn, 30 s, 9.3 KB · canvas, three.js, no errorsEndless runner · one pi turn, 98 s, 10.3 KB · canvas, three.js, no errorsSpace shooter · one pi turn, 92 s, 20.8 KB · canvas, three.js, no errors
Game
Wall · turns
Tool calls
Tokens in / out / cache
Decode tok/s mean (peak)
Prefill peak
Breakout
30 s · 4
write x2, edit
6,256 / 3,908 / 9,990
108 (156)
3,666
Endless runner
98 s · 11
bash x5, write, edit x3, read
9,160 / 6,202 / 62,833
66 (97)
3,975
Space shooter
92 s · 5
bash x2, write, edit
10,763 / 8,674 / 24,439
81 (116)
2,899
All three
220 s · 20
26,179 / 18,784 / 97,262
Token counts are pi's own per-turn usage (input, output, prompt-cache reads); throughput is from the server's per-request lines, so wall time includes pi's tool execution (the runner ran `node --check` and five shell commands). The one-turn prompt was the only instruction; no follow-up was needed.
Each game was requested in one prompt from pi (provider mlx, thinking on) against MLX-Serve 26.10.1 serving the Flash Next mixed 4-8 bit pack with MTP; every file arrived on the first turn and passed the load check (canvas present, three.js loaded, no page or console errors). The Starcraft-themed FPS comparison against TensorFold is below.
Starcraft-themed FPS, same prompt, two engines
pi builds a Terran-vs-Zerg FPS: MLX-Serve 26.10.1 vs TensorFold 0.6.0 (M5 Ultra)
one prompt, thinking on, headless-Chromium load check
MLX-Serve 26.10.1 · our Flash Next mixed 4-8 bit pack, MTP · 124 s, 11 turns, 40.3 KB · loads cleanTensorFold 0.6.0 · its Vontra Flash Next 4-bit MTP pack, a lower quant than ours · 230 s, 6 turns, 30.4 KB · loads clean
Arm
Wall · turns
Tool calls
Tokens in / out / cache
Decode tok/s per request
First write · final prompt
MLX-Serve 26.10.1 + our pack
124 s · 11
write, edit x2, bash x6, read x2
32,703 / 18,279 / 196,303
114-169, median 137
14.5k tok at 169 tok/s · 32.4k (32.1k cached)
TensorFold 0.6.0 + Vontra pack
230 s · 6
write, edit, bash x2, read
44,383 / 31,156 / 142,521
129-158, median 140
15.4k tok at 156 tok/s · 44.4k (43.6k cached)
Not apples to apples: TensorFold could not load our mixed 4-8 bit pack unmodified (it does not know the n-gram table's shard layout), so the third arm was skipped as agreed, and TensorFold ran the Vontra 4-bit pack it ships for, a lower quant with fewer bytes per token, which flatters its speed and costs it quality. Per-request decode landed in the same band (median 137 vs 140 tok/s) despite that handicap on our side; TensorFold's run wrote 70% more output tokens (longer thinking at effort medium and two 7k-token rewrites) and took 1.9x the wall time, mlx-serve's took more, shorter tool rounds. Both games pass the load check; playability is for a human to judge from the files.
Test results
Flash Next packs from bf16 with a 1.35 M-token imatrix (M5 Ultra)
scored on the 290 held-out positions all four packs answered, against the bf16 model's own logits; speed same session, ABBA, llmprobe
Pack
bpw (experts)
Weights
Top-1 all / agent / code / math / prose
KLD median all
KLD mean all
Decode tok/s (median, n)
Prefill
mixed 4-8, shipped (no imatrix)
4.68 (4.50)
70.1 GB
87.9 / 89.0 / 88.0 / 87.5 / 87.0
0.0066
0.158
166 (n=6)
5369
iQ-MLX-4.7bpw: mixed 4-8 + imatrix (same layout), published
4.68 (4.50)
70.1 GB
88.6 / 89.0 / 96.0 / 90.0 / 81.2
0.0052
0.126
157 (n=6)
5411
68 GB allocator + imatrix
4.93 (4.74)
73.9 GB
89.3 / 86.8 / 96.0 / 88.8 / 88.4
0.0048
0.183
156 (n=2)
5380
iQ-MLX-5.2bpw: 75 GB allocator + imatrix (local)
5.21 (5.04)
78.1 GB
91.7 / 91.2 / 96.0 / 92.5 / 88.4
0.0054
0.225
148 (n=2)
5039
Top-1 is the share of positions where the pack's greedy token equals bf16's. KLD is KL(bf16 || pack) over bf16's top-1024 tokens from 1024 returned logprobs, with tokens outside the pack's own top-1024 floored at its last value; the mean is dominated by a few flat prose positions where bf16's own top token sits under 5%, so the median is the steadier summary. One row is 0.3 points of top-1 overall and 1-2 points per slice.
The calibrated 4-8 pack is byte-for-byte the shipped layout (every expert 4-bit g64, non-expert 8-bit, embedding 4-bit), so it is a drop-in: 20% lower KLD mean and median, top-1 flat, decode within the MTP swing (shipped arms ran 141-173). The 5 bpw pack has the best top-1 on every slice but reads 8 GB more per pass: 11% slower decode, 6% slower prefill. The 68 GB allocator variant has no niche.
Imatrix: 1.35 M tokens (agent, code, prose, math incl. SWE-bench Lite) through the bf16 model, 298 min; allocation by measured reconstruction error per layer and role. The calibrated 4-8 layout is published as ddalcu/Qwen3.8-Flash-Next-MLX-Serve-iQ-MLX-4.7bpw (113 files, 107.3 GB, 2026-10-01 10:10 EDT); the 5.2bpw and 68 GB packs stay local.
Every cell measured so far
decode / prefill tok/s, context rungs as decode tok/s