VERIFIED UPDATE · 2 OCT 2026

MLX-Serve's new gate moves different frontiers.

The author published v26.10.1 on 1 October. The detailed same-box validation snapshot compares the shipped 26.9.6 with main commit 599dd05, alternating test arms with llmprobe. The release tag points to 02bee55, so these detailed cells are validation-main measurements, not a claim that the tagged binary was rerun for every cell.

QWEN3.8 27B · 4-BIT + DRAFTER242 TG146 → 242 tok/s · +66%; PP 1,689 → 1,917
FLASH-NEXT · MIXED 4/8-BIT5,295 PP3,518 → 5,295 tok/s · +51%; TG 158 → 154 is a tie
35B-A3B · CONTEXT UNKNOWN10,072 PP9,001 → 10,072 tok/s · TG 403 → 404

What the points mean

All three headlines come from an M5 Ultra with 256GB. The validation page does not disclose its GPU core count, the exact llmprobe prompt length, generation starting context, output length, per-cell cache state or reasoning mode. The paired chart highlights indicate the same source row, not a single API request with identical input context. The 35B value cannot be ranked against a known 32K or 64K cold prefill test. The release notes independently confirm the +66%, +51% and +12% directions.

The report's TensorFold Flash-Next row uses a different, lower-bit Vontra pack, so its speed is not a same-weights engine comparison. Its 27B 4/6/8-bit same-pack rows are more useful, but the 6/8-bit TensorFold arms also use their shipped drafter. We have kept engine, quantization and speculation beside each point.

A special gain, not a general decode gain

Bonsai-2 27B 2-bit gained 48% generation speed in a 2.4KB quoted Zig rewrite yielding 654 output tokens: 184.7 → 272.9 tok/s. Its novel-text run was nearly flat at 110 → 111.5. Both workloads stay distinct in the log. The release attributes the quoted-text gain to a new verify path. It also reports that a Qwen3.8 27B + drafter can now admit 16K context on a 32GB M1 Pro where the previous release refused requests.

Explore both charts