Ultra Flash-Next Q6 at 195K: 4,692 PP / 70.0 TG.
The original oMLX session reports Qwen3.8-Flash-Next-oQ6e-mtp on an M5 Ultra 80 GPU cores / 256GB. The 195K configured context bucket measures 4,692 prefill and 70.0 one-stream generation tokens/s. Its recipe explicitly has MTP off and DFlash off, despite “mtp” in the checkpoint name and inactive DFlash cache fields.
Same Ultra session, different context buckets
| Configured context | PP tok/s | One-stream TG tok/s | Original run |
|---|---|---|---|
| 1K | 1,755 | 85.3 | view |
| 8K | 4,649 | 79.8 | view |
| 64K | 4,962 | 77.3 | view |
| 128K | 4,827 | 70.6 | view |
| 195K | 4,692 | 70.0 | view |
Engine: oMLX 0.7.1.dev1 on macOS 27.0.1. The session's eight-stream batching total is 280.6 TG tok/s; its individual stream contexts and output lengths are not published. For the five context buckets, the exact tokenized prefill prompts, generation starting contexts, actual output lengths and cache reuse are unpublished. These buckets are not verified long-context recall tests.
Other new original results
- M5 Max 40c / 128GB Flash-Next, oMLX 0.7.1.dev1, 128K bucket: 2,299 PP / 54.0 TG. MTP depth 3 on, DFlash off. Differences from the prior 0.7.0 recipe prevent a controlled engine-speed claim.
- M5 Max 40c / 128GB Uncensored Flash-Next 5-bit, 16K: 1,927 / 61.5. A distinct derivative, with MTP and TurboQuant KV4 on, DFlash and SpecPrefill off.
- M5 Pro 20c / 64GB Swift-1.5 27B, 64K: 512 / 21.7. MTP on, DFlash and SpecPrefill off.
- M5 Max 40c / 48GB Abliterated 27B MXFP4, 128K: 546.5 / 21.9. MTP is off in the recipe even though the model name says MTP; 128K recall was not verified.
- M5 Ultra 64c / 256GB GLM-5.3-Flash-CYBER derivative, 32K: 1,894 / 54.9. MTP on, DFlash off. First published 9 October, kept separate from the base GLM model.
Independent tests with different conditions
An independent MLX-Serve Flash-Next mixed 4/8-bit context sweep, originally published 30 September, reports an M5 Ultra 256GB. At a requested 256K rung it records 4,114 PP / 114.4 one-stream TG tok/s; GPU core count is unpublished. The probe used a forced 192-token output and the author says its planted value was never used in answers, so the speed sweep does not verify 256K recall. It is a separate pack and engine from the Q6 result above.
The GLM llama.cpp author's report on an Ultra 80c / 256GB measures 764 PP tok/s at an exact 4,096-token llama-bench prompt and approximately 629 PP at an exact 58,806-token prompt, where recall was checked. A separate four-prompt generation test gives 61.8 TG tok/s with probabilistic MTP three drafts versus 39.7 without it. Exact generation starting context is unpublished. The 262K server setting is capacity; the author explicitly did not verify a real 262K prompt or recall at that depth.
Explore both charts ↗