Today's monitoring surfaced sources dated 2 October. Chart points stay on the sources' dates, not today's discovery date. The first is an oMLX community benchmark session for GLM-5.3-Flash-oQ4e on an M5 Ultra with 64 GPU cores and 256GB RAM. The second is merged oMLX PR #4052 with bracketed Qwen Flash-Next kernel tests.
QWEN · THE GRILL218.3 TGNamed workload · not a generic rate
GLM: a different 64-core recipe
The session reports configured 1K at 1,364 PP / 50.6 TG, 4K at 1,987 / 41.1, and 64K at 1,918 / 51.6 tok/s. It runs oMLX 0.7.0 on macOS 27.0.1 with DFlash2 adaptive verification, MTP off and 4-bit TurboQuant KV. The DFlash in-memory cache is configured, but a per-run cache hit is not reported; its SSD cache is off. The actual tokenized prompt, generation starting context, output length and reasoning output are unpublished. This oQ4e checkpoint is separate from the earlier mixed 4/8-bit 64-core result and the 80-core MLX-4bit result, so their speeds do not form a controlled A/B.
Qwen: merged decode kernels on 5-bit weights
With fresh servers and prompt cache off, PR #4052 reports 112.7–113.6 → 120.0 TG with MTP off, then 178.5–179.1 → 189.2–189.4 TG with Lightning MTP on, over four coding prompts of 400 generated tokens each. The same source's The Grill workload rises from 201.9–204.0 to 218.3 TG. The author reports byte-identical greedy output for the four prompts and a merged PR. Its M5 Ultra trace says 80 GPU cores; RAM size, exact generation starting context and power were not published. The 24K The Grill prefill result, 4,926 PP, was approximately flat, as was four/eight-stream aggregate output. These decode results use oQ5e 5-bit weights and should not be conflated with the open uniform 4-bit MTP equivalence report.