SOURCE UPDATE · 9 OCT 2026

ANE prefill on Max. A modest Ultra kernel gain.

Two original oMLX runs dated 8 October report Qwen3.8-27B-MLX-Serve-4bit on an M5 Max with 40 GPU cores and 128GB. The actual engine is oMLX 0.7.0 on macOS 27.2; the checkpoint name does not turn this into an mlx-serve HTTP load test. Qwen ANE Prefill and GDN were on. MTP, DFlash, SpecPrefill and TurboQuant KV were off.

MAX 40C · CONFIGURED 8K681.8PP tok/s; TG 32.6
MAX 40C · CONFIGURED 32K595.4PP tok/s; TG 29.4
ULTRA 256GB · SHORT PROMPTS+2.5%Separate mlx-serve TG A/B

One Max session, two context settings

Configured bucketPP tok/sOne-stream TG tok/sTTFTSource
8K681.832.612,016 msoriginal run
32K595.429.455,036 msoriginal run

8K and 32K are configured context buckets, not published exact tokenized prompt lengths or generation starting depths. Actual output length and per-run prompt-cache reuse are unknown. The daily log keeps the two buckets as separate records, and labels this hardware as M5 Max, not M5 Ultra. No Ultra extrapolation was added: ANE prefill scaling to a different chip has not been measured here.

Merged mlx-serve kernel A/B

PR #732, first published 5 October and merged 6 October, changes the G17 MoE down+reduce kernel from 8 to 16 lanes. On an M5 Ultra 256GB using Flash-Next 4.7bpw (4-bit experts, 8-bit elsewhere), its interleaved, one-client, 12-short-prompt A/B pooled 169.5 to 173.8 output tok/s, a 2.5% TG change. MTP and PLE-GPU were on. A prefix-cache capacity of 4GB was configured, but cache reuse was not reported. Prompt inputs were at most 400 tokens; exact TG starting context, generated length, reasoning and DFlash state were not disclosed. The authors note about 1% output-token-count drift and roughly 2% baseline variation across windows. This PR provides no paired PP measurement and does not establish a general model speed record.

Explore both charts