SOURCE UPDATE · 10 OCT 2026

Ultra Ornith at 32K: 7,405 PP / 178.5 TG.

The original 9 October oMLX result is an M5 Ultra with 80 GPU cores and 96GB running the distinct Ornith-1.5-35B-A3B-oQ4e-mtp derivative. It is not a Qwen3.6 35B-A3B result or a 256GB Ultra result. The same session reports seven configured context buckets from 1K to 128K.

ULTRA 80C · ORNITH · 32K7,405PP tok/s · TG 178.5
MAX 40C · FLASH-NEXT · 128K1,745PP tok/s · TG 55.8
PRO 20C · 27B MTP · 16K670.7PP tok/s · TG 48.6

Ultra 80-core / 96GB Ornith 1.5

Configured contextPP tok/sOne-stream TG tok/sPeak memoryOriginal run
1K6,312150.021.4GBview
4K9,793155.422.0GBview
8K9,095161.622.1GBview
16K8,826173.622.3GBview
32K7,405178.522.7GBview
64K5,739129.123.4GBview
128K3,864114.624.9GBview

Engine: oMLX 0.7.0 on macOS 27.0.1, Code (Python) workload. MTP adaptive depth 3 is on; DFlash, SpecPrefill, TurboQuant KV and ANE prefill are off. Thinking is enabled but the configured 8,192-token thinking budget is disabled. Exact tokenized prompts, TG starting contexts, generated lengths and cache reuse are unpublished. Each bucket is its own pair of metric records in the tracker; the two charts link the matching source on hover.

Max 40-core / 128GB Flash-Next to 195K

The separate M5 Max session uses Qwen3.8-Flash-Next-oQ4e-mtp, oMLX 0.7.0 and macOS 26.6.2. Its 128K result is 1,745 PP / 55.8 TG tok/s; the 195K configured bucket is 1,240 / 46.1. At 4K, the run is 1,935 / 79.9. The site records all seven context buckets separately and labels them Max comparison runs. MTP adaptive depth 3 is on. DFlash is off, even though dormant DFlash cache settings appear in the recipe. The 40,000-token thinking budget and 80,000 output maximum are configured limits, not generated token counts; actual output and cache hits are unknown. No Max-to-Ultra extrapolation was added.

Pro 20-core / 64GB 27B recipes

A 27B MLX-4bit 1K run reports 459.5 PP / 17.5 TG, with thinking budget and TurboQuant KV 8-bit on, MTP off. Its 4K run is 477.8 / 17.2. A separate oQ4e-mtp 4K run is 638 / 24.5, while its 16K run is 670.7 / 48.6. That second recipe has A8 activation and MTP adaptive depth 6 on, but its thinking budget is off. Checkpoint, OS, context and thinking settings differ, so these are not a controlled MTP on/off test.

Explore both charts