VERIFIED UPDATE · 1 OCT 2026

Stable oMLX, new Ultra tests.

oMLX 0.7.0 was released on 30 September. This release includes the Qwen3.8 and GLM-5.3 M5 kernel work and GLM DFlash2 support. The release notes themselves are a software milestone, not a new speed measurement. Two later community benchmark sessions on a 64-core M5 Ultra with 256GB were uploaded late on 30 September GMT, which is 1 October in Taiwan.

FLASH-NEXT Q4E · CONFIGURED 1K1,440 / 89.7PP / one-stream TG tok/s · stable 0.7.0
GLM MIXED 4/8 · CONFIGURED 64K1,909 / 33.7PP / one-stream TG tok/s · DFlash2
QWEN Q4E · 100K+ PROMPT228.3eight-stream aggregate TG · draft PR, MTP off

Read each workload on its own terms

The Flash-Next oQ4e + MTP run reports 1,440 prefill and 89.7 single-stream generation tok/s at configured 1K context. Its separate four-stream batch table shows 172.1 tok/s across the device. The GLM-5.3 mixed-4/8-bit DFlash2 session reports 1,939 / 33.3 at configured 32K, 1,909 / 33.7 at 64K, and 1,873 / 29.6 at 128K. Both pages give context settings, but not exact tokenized prompt lengths, generated output lengths, or per-run prefix-cache hits. The 1K, 32K, 64K and 128K points remain separate. Neither session replaces the site's default payback recipe.

For long prompts, draft PR #4070 tested eight sessions with a 100K-token shared prompt plus 2K unshared tokens each, generating 768 tokens per session. With its new batched QSA paths, device-wide generation went from 93.3 to 228.3 tok/s with MTP off; with MTP on, 84.7 to 150.9. The author disclosed an M5 Ultra with 256GB but not GPU core count. These are aggregate draft-build results, not single-request speeds or a stable-release baseline.

Two changes that affect how to read numbers

Merged GLM DFlash2 PR #3979 separately reported 62.4, 47.3 and 44.0 single-stream tok/s after 200, 1,024 and 4,096 prompt tokens on an 80-core Ultra. It used GLM oQ4e and a test build with local M5 kernel patches; it is not the same run as the new 64-core mixed-4/8-bit session. PR #4087 reduced oQ8e loading memory from 235.1 to 181.5 GiB on a 256GB Ultra; no new PP/TG rate was supplied.

Merged usage fix #4065 stopped counting cached prefix tokens as freshly prefilled tokens in API rates. A documented cache-hit request had displayed 23,448 PP tok/s even though fresh prefill on that machine was around 1,700. Our chart labels unknown cache states as unknown; it does not treat such API rates as cold prefill.

Explore both charts