SOURCE UPDATE · 8 OCT 2026

Fast at 8K. Still measured at 195K.

We checked a 7 October oMLX session on an M5 Ultra with 80 GPU cores and 256GB. It used Qwen3.8-Flash-Next oQ4e 4-bit on oMLX 0.7.0, with Lightning MTP fixed depth 2 on. DFlash, SpecPrefill and TurboQuant KV were off. This is one community-uploaded session, not independent replication.

8K BUCKET · ONE STREAM5,627Prefill tok/s
8K BUCKET · ONE STREAM147.7Generation tok/s
195K LABEL · ONE STREAM90.3Generation tok/s; PP 5,226

The original session, context by context

Configured bucketPP tok/sTG tok/sSource
1K2,890144.1run
4K4,905132.7run
8K5,627147.7run
64K5,553127.5run
128K5,379106.2run
195K5,22690.3run

The page's context values are configured settings, not published exact tokenized prompt or TG starting lengths. The 195K label has a 200000 configured context field. Generated length and per-run prompt-cache reuse are unknown. DFlash memory/SSD limits appearing in the recipe do not mean DFlash ran; dflash_enabled=false.

A separate test with exact input counts

An independent author published a 6 October Mac versus two-DGX comparison. The Mac was an M5 Ultra 256GB with an optimized 4-bit MLX Flash-Next build; its GPU core count and exact engine/kernel version were not stated. Each case used the same prompt and temperature on both sides, one warm-up and three timed runs. The author's Mac table reports 1,987 input tokens: 2,932.9 PP / 140.0 TG; 8,489: 2,931.0 / 109.3; and 14,614: 3,260.1 / 129.3 tok/s. The last two Mac completions produced about 44 tokens despite requesting 128. MTP, DFlash, reasoning and timed-run cache reuse were not specified. These points have their own series; they are not joined to oMLX's different recipe.

Explore both charts