APPLE SILICON · OPEN BENCHMARK WATCH

How fast is the
M5 Ultra getting?

A living record of local AI inference—prefill, token generation, the kernels behind each move, and what the numbers could mean for selling tokens.

Latest sourced event: 28 Sep 2026 · Public log through 29 Sep 2026

TWO SEPARATE TESTS / 28 SEP 2026● ● ●
PREFILL · 32K4,759input tok/sFlash-Next · oQ6e · 80c · cache unknown
GENERATION · CONTEXT UNKNOWN158–165output tok/sFresh prompt · 1 stream · output unknown
32K PREFILL EVOLUTION3,579 → 4,759
27 SEP28 SEP
NO MATCHED PP/TG PAIR CLAIMEDVIEW SOURCE ↗
01Two metrics, never one vague speed number
02Context, quantization and concurrency stay visible
03Every data point links to original evidence
01 / PERFORMANCE

The progress, day by day.

The time axis runs earliest to latest. Every point still shows its PP, TG and test conditions.

FIRST VERIFIED ULTRA DAY21 SEP 2026
◎
01 / PREFILL

Read the prompt

INPUT TOK/S

Cold prompt processing. Higher is faster under the same context and recipe.

02 / TOKEN GENERATION

Write the answer

OUTPUT TOK/S

One stream and aggregate multi-stream results are labelled separately.

OTHER MACS / WHAT-IF

Inco Splash, M5 Pro & M5 Max.

Measured source-Mac throughput and a separate, unverified M5 Ultra scaling range. Source context stays attached to every value.

倍率依據 / Scaling basis

M5 Pro 16c → Ultra 80c: PP ×2–4, TG ×1.5–2.5. M5 Max 40c → Ultra 80c: PP ×1.2–1.8, TG ×1.2–1.7. Apple lists 307 / 614 / 1,200 GB/s memory bandwidth and 16 / 40 / 80 GPU cores for these specific variants. These broad factors allow for less than linear software scaling; they are scenarios, not cross-device benchmark findings. Each projection preserves its own context and engine.

Apple Pro / Max specs ↗ · Apple Ultra specs ↗

Official Splash context curve · M5 Pro 16c / 48GB

Model / metricShort8K bucket16K bucket32K bucket
35B PP · uncached—2,5752,3732,011
35B TG · 1 stream210156149143
27B PP · uncached—398395363
27B TG · 1 stream74555554

SPEED-Bench coding prompts, reasoning on (27B medium), DFlash 2, HTTP, output cap 1,024; actual output lengths and TG cache state unpublished. Four short-stream aggregate: 35B 357; 27B 170. Cached 32K replay TTFT is a separate test. Inco official results ↗

Weinbach prototype v13–v17 · M5 Ultra

Metric / contextv13v14v15v16v17
PP · 2K prompt4,0144,0974,1964,2244,180–4,300
TG · 10 fresh, starting context unknown51.5102.5118.9115.9165.7
TG · 2K coding, 1 stream62.075.0124.2*85.9179.4
TG · 2 streams aggregate———153.2138.2
TG · 4 streams aggregate———220.3159.1

All versions appeared in one post on 24 Sep (Taipei). Author confirmed MTP and said v17 batch is broken. Exact quantization, implementation, fresh-prompt context, output length, cache/reasoning and batch context were not published in the table. *v15 has an unexplained source footnote. Original table ↗ · Author reply ↗

02 / CHANGELOG

What actually changed.

Newest sourced results first; the oldest records are at the bottom. Each model and patch links to its source.

Share latest update ↗
03 / TOKEN ECONOMICS

What could this earn?

A transparent capacity scenario, not observed demand or a promise of profit.

SCENARIO / NOT REVENUE
SELECT A RECIPE

Pricing & operating assumptions+
AT YOUR ASSUMPTIONSMEASURED PAIR
Monthly operating cash——
Gross / day—
Gross / month—
Simple payback—
UTILIZATION SENSITIVITYNET / MONTH

Capacity assumes prefill and decoding consume the device sequentially. Real sales also depend on queueing, latency, caching, quality, uptime, distribution and demand.

04 / CAPACITY WAITLIST

Find or offer Mac inference capacity.

Join an interest list. This is not a live rental marketplace or a booking.

Availability, pricing, model rights and service terms still need separate confirmation. We will only contact you about this waitlist and potential matches.

We store your email, selected interest, optional details and consent time in our Cloudflare database. Email ownership is not verified yet; no messages are sent today. After signup you receive a private withdrawal link—save it. Your information is not shown publicly.

05 / READ THE DATA

The conditions matter.

01

Prefill ≠ generation

Prefill reads input. Generation writes output. A fast result in one stage cannot stand in for the other.

02

Keep recipes together

Compare model, chip, quantization, context, cache state, prompt and concurrency. Different recipes appear as separate points.

03

Commercial use needs its own check

Flash-Next's model license requires separate permission for Model as a Service. Price snapshots and theoretical capacity are not contracted revenue.

MODEL LICENSE ↗
CAPACITY MODELoutput tok/s = 1 / (input:output ratio / PP + 1 / TG)

A conservative sequential-time approximation. Multi-user serving should use measured aggregate throughput with a latency target.