Fast at 8K. Still measured at 195K.
We checked a 7 October oMLX session on an M5 Ultra with 80 GPU cores and 256GB. It used Qwen3.8-Flash-Next oQ4e 4-bit on oMLX 0.7.0, with Lightning MTP fixed depth 2 on. DFlash, SpecPrefill and TurboQuant KV were off. This is one community-uploaded session, not independent replication.
The original session, context by context
| Configured bucket | PP tok/s | TG tok/s | Source |
|---|---|---|---|
| 1K | 2,890 | 144.1 | run |
| 4K | 4,905 | 132.7 | run |
| 8K | 5,627 | 147.7 | run |
| 64K | 5,553 | 127.5 | run |
| 128K | 5,379 | 106.2 | run |
| 195K | 5,226 | 90.3 | run |
The page's context values are configured settings, not published exact tokenized prompt or TG starting lengths. The 195K label has a 200000 configured context field. Generated length and per-run prompt-cache reuse are unknown. DFlash memory/SSD limits appearing in the recipe do not mean DFlash ran; dflash_enabled=false.
A separate test with exact input counts
An independent author published a 6 October Mac versus two-DGX comparison. The Mac was an M5 Ultra 256GB with an optimized 4-bit MLX Flash-Next build; its GPU core count and exact engine/kernel version were not stated. Each case used the same prompt and temperature on both sides, one warm-up and three timed runs. The author's Mac table reports 1,987 input tokens: 2,932.9 PP / 140.0 TG; 8,489: 2,931.0 / 109.3; and 14,614: 3,260.1 / 129.3 tok/s. The last two Mac completions produced about 44 tokens despite requesting 128. MTP, DFlash, reasoning and timed-run cache reuse were not specified. These points have their own series; they are not joined to oMLX's different recipe.
Explore both charts