26.10.1 Release Gate
mlx-serve · release validation · interim

26.10.1 Release Gate

MLX-Serve 26.10.1 (main 599dd05, the release candidate) against the shipped MLX-Serve 26.9.6, same session, arms alternating, llmprobe measuring; TensorFold beside them where the same pack was measured on the same box. Decode and prefill in tokens per second. Snapshot taken 2026-10-01 11:13 EDT while the runs continue on four Macs.

MLX-Serve 26.9.6MLX-Serve 26.10.1TensorFold (0.6.0 on every M5 Ultra row, 0.5.0 elsewhere; same box, same llmprobe method; the 6/8-bit rows run its drafter)
Agent build check

Three three.js games by pi on Flash Next (MLX-Serve 26.10.1, M5 Ultra)

one prompt each, single-file index.html, loaded in headless Chromium
Breakout built by pi + Flash Next, first frame in headless Chromium
Breakout · one pi turn, 30 s, 9.3 KB · canvas, three.js, no errors
Endless runner built by pi + Flash Next, first frame in headless Chromium
Endless runner · one pi turn, 98 s, 10.3 KB · canvas, three.js, no errors
Space shooter built by pi + Flash Next, first frame in headless Chromium
Space shooter · one pi turn, 92 s, 20.8 KB · canvas, three.js, no errors
GameWall · turnsTool callsTokens in / out / cacheDecode tok/s mean (peak)Prefill peak
Breakout30 s · 4write x2, edit6,256 / 3,908 / 9,990108 (156)3,666
Endless runner98 s · 11bash x5, write, edit x3, read9,160 / 6,202 / 62,83366 (97)3,975
Space shooter92 s · 5bash x2, write, edit10,763 / 8,674 / 24,43981 (116)2,899
All three220 s · 2026,179 / 18,784 / 97,262
Starcraft-themed FPS, same prompt, two engines

pi builds a Terran-vs-Zerg FPS: MLX-Serve 26.10.1 vs TensorFold 0.6.0 (M5 Ultra)

one prompt, thinking on, headless-Chromium load check
FPS built on mlx-serve, title screen
MLX-Serve 26.10.1 · our Flash Next mixed 4-8 bit pack, MTP · 124 s, 11 turns, 40.3 KB · loads clean
FPS built on TensorFold, title screen
TensorFold 0.6.0 · its Vontra Flash Next 4-bit MTP pack, a lower quant than ours · 230 s, 6 turns, 30.4 KB · loads clean
ArmWall · turnsTool callsTokens in / out / cacheDecode tok/s per requestFirst write · final prompt
MLX-Serve 26.10.1 + our pack124 s · 11write, edit x2, bash x6, read x232,703 / 18,279 / 196,303114-169, median 13714.5k tok at 169 tok/s · 32.4k (32.1k cached)
TensorFold 0.6.0 + Vontra pack230 s · 6write, edit, bash x2, read44,383 / 31,156 / 142,521129-158, median 14015.4k tok at 156 tok/s · 44.4k (43.6k cached)
Test results

Flash Next packs from bf16 with a 1.35 M-token imatrix (M5 Ultra)

scored on the 290 held-out positions all four packs answered, against the bf16 model's own logits; speed same session, ABBA, llmprobe
Packbpw (experts)WeightsTop-1 all / agent / code / math / proseKLD median allKLD mean allDecode tok/s (median, n)Prefill
mixed 4-8, shipped (no imatrix)4.68 (4.50)70.1 GB87.9 / 89.0 / 88.0 / 87.5 / 87.00.00660.158166 (n=6)5369
iQ-MLX-4.7bpw: mixed 4-8 + imatrix (same layout), published4.68 (4.50)70.1 GB88.6 / 89.0 / 96.0 / 90.0 / 81.20.00520.126157 (n=6)5411
68 GB allocator + imatrix4.93 (4.74)73.9 GB89.3 / 86.8 / 96.0 / 88.8 / 88.40.00480.183156 (n=2)5380
iQ-MLX-5.2bpw: 75 GB allocator + imatrix (local)5.21 (5.04)78.1 GB91.7 / 91.2 / 96.0 / 92.5 / 88.40.00540.225148 (n=2)5039

Every cell measured so far

decode / prefill tok/s, context rungs as decode tok/s