VERIFIED SOURCE UPDATE · 3 OCT 2026

Long context moves forward. The conditions still matter.

Today's monitoring uncovered benchmark sessions dated 1–2 October. The source dates, not today's discovery date, position the new chart points. The oMLX 0.7.0 Qwen session uses a 64-core, 256GB M5 Ultra, a 4-bit oQ4e Flash-Next checkpoint and MTP off.

QWEN · CONFIGURED 195K4,547 PP96.1 single-stream TG · MTP off
QWEN · 8 STREAMS326.4 TGDevice-wide total · batch context not separately published
GLM · CONFIGURED 128K75.5 TG2,199 PP · 80 cores · DFlash2 on

Two distinct serving recipes

The Qwen session reports 4,827 PP / 98.6 TG at configured 64K, 4,685 / 96.7 at 128K, and 4,547 / 96.1 at 195K. Its batching table reaches 326.4 output tokens/s across eight streams. Batch B1 matches the session's 1K single-stream value, but the source does not independently list each batch request's actual context or output length. This is not eight 195K requests and does not replace a payback recipe.

A separate GLM-5.3-Flash 4-bit session on an 80-core M5 Ultra uses DFlash2 on, MTP off: configured 32K 2,275 PP / 77.0 TG, 64K 2,249 / 65.0, 128K 2,199 / 75.5. These are speculative generation rates, not target-only decode. Another 80-core Flash-Next Uncensored oQ6e session reached 4,671 PP / 139.2 TG at configured 8K and 4,465 / 101.0 at 128K, using MTP and 8-bit TurboQuant KV. It is a derivative checkpoint and is displayed separately from the original model weights.

Accuracy and admission notes

In open oMLX issue #4209, a reporter found different greedy outputs with MTP on versus off for four prompts on one Vontra uniform 4-bit group-size-32 checkpoint. Repeated MTP-on runs were consistent; a 5-bit control stayed byte-identical. The cause is unconfirmed and the report does not show that every 4-bit model diverges or that answer quality fell. The issue's roughly 155 versus 48.6 TG comparison must not be described as an exact-output speedup.

Open TensorFold PR #245 corrects a GLM memory estimator that misread a dense-to-sparse transition as continuing growth. On a 256GB M5 Ultra, it changes the estimated 200K prefill workspace from 249.64 to 6.81 GiB versus an 8.58GiB measured peak, allowing two 113K prompts alongside a live stream. Two 33K requests still took 62.7 seconds versus 61.7 before: the improvement is admission and fairness, not raw PP/TG.

Explore both charts