DAILY UPDATE · 28 SEP 2026
Kernel work moved the prefill frontier.
The newest pinned Qwen3.8 Flash-Next recipe on M5 Ultra 80c / 256GB reports a 32K prefill of 4,759 tok/s. Fresh generation reaches 158–165 tok/s in a separate test. Do not treat them as a measured 32K API request pair.
What changed
On 27 Sep, the Q6 route recorded 3,579 PP with the rc1 recipe and 4,733 PP after the upstream prefill stack. On 28 Sep, a decode-focused version fell to 4,615 PP while improving fresh generation. Subsequent kernel work restored PP to 4,734, then 4,759. The curve includes this regression rather than smoothing it away.
- Qwen Q6: 4,759 PP at 32K, 158–165 fresh TG, and 197 aggregate TG at eight streams are distinct measurements in the pinned recipe.
- TensorFold Q4: 327 aggregate TG at eight streams and 2,826 PP at 32K. Its reported KLD differs substantially from the Q6 recipe, so speed is not a quality-neutral comparison.
- GLM-5.3 Flash: the pinned build reports 2,233 PP at 32K and 79–81 fresh TG in separate tests.
The commercial calculator on the main site is a scenario. Qwen3.8-Flash-Next's license requires separate permission for Model as a Service.
Explore the full timeline