VERIFIED BACKFILL · 29 SEP 2026

The long-context launch-day results.

Today's update adds measurements from Federico Viticci's MacStories review, published 21 September. These are older experiments added to the public chart today, not new tests performed on 29 September.

FLASH-NEXT · COLD PREFILL · 261,856 INPUT2,544input tokens/s · short retrieval answer
FLASH-NEXT · 512-TOKEN ANSWER · 261,597 INPUT74.7output tokens/s · one stream, no cache
QWEN3.8 27B · 6,091 INPUT1,701 / 48PP / TG tokens/s · one 221-token answer

Keep the workloads separate

For Flash-Next oQ4e on the 80-core M5 Ultra, the author's cold prefill tests recorded 2,732, 2,654 and 2,544 tokens/s at 65,235, 130,781 and 261,856 input tokens. Those trials asked for short retrieval answers. A separate single-run curve asked for a 512-token answer after 3,578, 15,867, 65,005, 130,522 and 261,597 input tokens; generation was 90.7, 87.8, 83.8, 60.6 and 74.7 tokens/s respectively. The chart keeps each precise input length and records these as different points.

At 6,091 input tokens, the review's Qwen3.8 27B oQ4e Mac run measured 1,701 prefill and 48 generation tokens/s over a 221-token answer. The author states that test used thinking off and no multi-token prediction. At 63,445, 128,983 and 260,037 input tokens, its generation rates were 38.9, 32.4 and 24.3 tokens/s. The exact output token counts for those long-context trials were not published.

New M5 Max comparisons from the Spark lead list

We checked Spark's leads against each original source. On a 40-core M5 Max with 64GB, the oMLX 0.7.0rc1 27B session reports 4K configured context at 974.3 PP / 35.6 TG tok/s; its 32K run reports 929.3 / 33.1. The full 1K–32K curve appears in the comparison section. MTP and DFlash were off; actual prompt and output token counts were not disclosed. An older 22 September oMLX 0.6.4 MTP-on run shows 815.9 PP / 56.4 TG at 1K configured context. Its 1,255 figure is TTFT in milliseconds, not output tokens.

The independent Splish author's pinned results compare the unofficial Splash 1.1.0 fork against stock Splash on a 40-core M5 Max with 128GB. For Qwen3.8-27B at 8K, one-stream TG rises from 74 to 95 tok/s; at 128K, from 54 to 63. Cold prefill at 2K falls from 1,019 to 923; the author says prefill kernels did not change. A separate long-reasoning test reports 177.7 one-stream and 391.8 four-stream aggregate TG, with starting context unpublished. These numbers use DFlash2 drafting and are not comparable to the MTP-on or MTP-off oMLX runs. All Max figures remain separate from Ultra chart points and default payback scenarios.

Explore both charts