SOURCE UPDATE · 11 OCT 2026

Ultra Flash-Next Q6 at 195K: 4,692 PP / 70.0 TG.

The original oMLX session reports Qwen3.8-Flash-Next-oQ6e-mtp on an M5 Ultra 80 GPU cores / 256GB. The 195K configured context bucket measures 4,692 prefill and 70.0 one-stream generation tokens/s. Its recipe explicitly has MTP off and DFlash off, despite “mtp” in the checkpoint name and inactive DFlash cache fields.

ULTRA 80C / 256GB · Q6 · 195K4,692PP tok/s · configured bucket
SAME RESULT · ONE STREAM70.0TG tok/s · exact start unknown
SEPARATE EIGHT-STREAM TEST280.6whole-device TG tok/s, not per stream

Same Ultra session, different context buckets

Configured contextPP tok/sOne-stream TG tok/sOriginal run
1K1,75585.3view
8K4,64979.8view
64K4,96277.3view
128K4,82770.6view
195K4,69270.0view

Engine: oMLX 0.7.1.dev1 on macOS 27.0.1. The session's eight-stream batching total is 280.6 TG tok/s; its individual stream contexts and output lengths are not published. For the five context buckets, the exact tokenized prefill prompts, generation starting contexts, actual output lengths and cache reuse are unpublished. These buckets are not verified long-context recall tests.

Other new original results

Independent tests with different conditions

An independent MLX-Serve Flash-Next mixed 4/8-bit context sweep, originally published 30 September, reports an M5 Ultra 256GB. At a requested 256K rung it records 4,114 PP / 114.4 one-stream TG tok/s; GPU core count is unpublished. The probe used a forced 192-token output and the author says its planted value was never used in answers, so the speed sweep does not verify 256K recall. It is a separate pack and engine from the Q6 result above.

The GLM llama.cpp author's report on an Ultra 80c / 256GB measures 764 PP tok/s at an exact 4,096-token llama-bench prompt and approximately 629 PP at an exact 58,806-token prompt, where recall was checked. A separate four-prompt generation test gives 61.8 TG tok/s with probabilistic MTP three drafts versus 39.7 without it. Exact generation starting context is unpublished. The 262K server setting is capacity; the author explicitly did not verify a real 262K prompt or recall at that depth.

Explore both charts ↗