Read the settings behind the speed.
Today we verified three 6 October oMLX sessions and corrected the status of PR #4248. These are community or developer measurements, not independent replications. Chart dates follow the original source, not our publication date.
What each 6 October run actually used
The Ultra 64-core / 256GB run used a 6-bit Uncensored Flash-Next derivative. It reports 2,099 PP and 91.9 single-stream TG for a configured 1K Code (Python) bucket, plus 263.2 TG total at batch size four. Although the checkpoint name ends in -mtp, its recipe has mtp_enabled=false. DFlash and SpecPrefill are also off. The listed DFlash cache limits are inactive settings, not measured cache use.
The Max 40-core / 128GB run used Qwen3.6-35B-A3B oQ4 with A8 activation and Lightning MTP depth 3 enabled. It reports 6,304 PP and 187.4 one-stream TG in a configured 4K bucket. Another same-day run reports 6,250 / 177.1. SpecPrefill and DFlash are off in the recipes. This Max result is shown as an other-Mac comparison, not an Ultra measurement or a new payback default.
The Pro 20-core / 48GB REAP-288 derivative reports 268 PP and 9.4 TG in a configured 4K bucket. Expert offload is on; MTP, DFlash and SpecPrefill are off. The model is pruned and is kept separate from the unpruned Flash-Next series.
All oMLX context labels above are configured buckets. The pages do not publish each run's exact tokenized prompt, TG starting context, generated length or prompt-cache reuse. Separate PP and TG conditions remain attached to the chart points.
Two source corrections
oMLX #4248 was merged on 5 October as commit 84e3b43. The previously recorded 5,469 PP at configured 32K and 218.9 TG with MTP are developer A/B measurements of this merged code path; they are not a tagged oMLX 0.7.0 result. A surfaced GLM 128K Ultra result shows 1,843 PP / 74.8 TG, but its source date is 3 October. We backfilled it on that date instead of presenting it as a new 7 October test.