DAILY UPDATE · 28 SEP 2026

Kernel work moved the prefill frontier.

The newest pinned Qwen3.8 Flash-Next recipe on M5 Ultra 80c / 256GB reports a 32K prefill of 4,759 tok/s. Fresh generation reaches 158–165 tok/s in a separate test. Do not treat them as a measured 32K API request pair.

FLASH-NEXT Q6 · 32K PREFILL4,759input tok/s · oMLX main plus #4030
FLASH-NEXT Q6 · FRESH GENERATION158–165output tok/s · separate single-stream test
TENSORFOLD Q4 · 8 STREAMS327aggregate tok/s · different quantization quality

What changed

On 27 Sep, the Q6 route recorded 3,579 PP with the rc1 recipe and 4,733 PP after the upstream prefill stack. On 28 Sep, a decode-focused version fell to 4,615 PP while improving fresh generation. Subsequent kernel work restored PP to 4,734, then 4,759. The curve includes this regression rather than smoothing it away.

The commercial calculator on the main site is a scenario. Qwen3.8-Flash-Next's license requires separate permission for Model as a Service.

Explore the full timeline