SOURCE UPDATE · 6 OCT 2026

One checkpoint, separate tests.

The 6 October check found a 5 October mlx-serve commit and two dated oMLX community sessions. The chart uses their 5 October source date. These are developer/community reports, not independent replications.

ULTRA · OQ4E · 10K PP TEST5,383input tok/s · mlx-serve development commit
ULTRA · OQ4E · MTP DECODE172output tok/s · one stream · context unknown
M5 MAX · 27B · 64K BUCKET775 / 30.1PP / TG · 40 cores · derivative model

mlx-serve reaches the oMLX pack's fast paths

The author says the Jundot Qwen3.8-Flash-Next oQ4e checkpoint previously loaded but missed several tuned mixed-width paths. Commit cd07bf8 adds hyper-connection, MoE/QSA and MTP verify handling plus embedded n-gram table support. On M5 Ultra, the author reports a 10K prefill test 5,075 → 5,383 PP, a different single-stream MTP decode test 124 → 172 TG, 177 TG with --ple-gpu, and 144 → 186 TG across four requests. Four-stream output is device-wide, not per user. The 10K label belongs to the PP test; the TG starting context and generated length were not published. This is a development commit under the unreleased v26.10.2 changelog, not a tagged release. The tests disclose no matched PP/TG workload suitable for a default revenue estimate.

Two 27B comparison sessions

An M5 Max 40-core / 64GB result for a Qwen3.8-27B-Abliterated 4-bit derivative reports 775.1 PP / 30.1 TG in its configured 64K Code (Python) bucket. Its settings have enable_thinking=true but thinking_budget_enabled=false: the configured 8,192-token budget was not active. MTP depth 3 and TurboQuant KV 8-bit are on; DFlash is off. Keep this derivative separate from the base 27B checkpoint.

An M5 Pro 20-core / 64GB oQ8e session reports configured 1K: 395.9 PP / 20.4 TG and 4K: 412.6 PP / 21.5 TG. A8 activation and adaptive Lightning MTP depth 3 are enabled. DFlash is off, despite its unused memory-cache limit appearing in the recipe. Both sessions omit exact tokenized prompt length, TG starting context, generated length and per-run cache reuse. Neither is an Ultra measurement; no Ultra multiplier is applied to these new points.

Explore both charts