APPLE SILICON · OPEN BENCHMARK WATCH

How fast is the
M5 Ultra getting?

A living record of local AI inference—prefill, token generation, the kernels behind each move, and what the numbers could mean for selling tokens.

Latest sourced event: 28 Sep 2026 · Public log through 29 Sep 2026

LATEST VERIFIED / 28 SEP 2026● ● ●
PREFILL · 32K4,759input tok/sFlash-Next · oQ6e · 80c
FRESH GENERATION158–165output tok/sSeparate test · 1 stream
32K PREFILL EVOLUTION3,579 → 4,759
27 SEP28 SEP
MEASURED • DIFFERENT WORKLOADSVIEW SOURCE ↗
01Two metrics, never one vague speed number
02Context, quantization and concurrency stay visible
03Every data point links to original evidence
01 / PERFORMANCE

The progress, day by day.

First documented dates, in order. Lines connect only comparable recipes. A drop is part of the story.

FIRST VERIFIED DAY23 SEP 2026
PrefillGeneration
◎
INPUT / PREFILL

Read the prompt

tok/s

Cold prompt processing. Higher is faster, under the same context and recipe.

OUTPUT / TOKEN GENERATION

Write the answer

tok/s

One stream and aggregate multi-stream results are labelled separately.

02 / CHANGELOG

What actually changed.

From the first verified day forward—model by model, patch by patch.

Share latest update ↗
03 / TOKEN ECONOMICS

What could this earn?

A transparent capacity scenario, not observed demand or a promise of profit.

SCENARIO / NOT REVENUE
SELECT A RECIPE

Pricing & operating assumptions+
AT YOUR ASSUMPTIONSMEASURED PAIR
Monthly operating cash——
Gross / day—
Gross / month—
Simple payback—
UTILIZATION SENSITIVITYNET / MONTH

Capacity assumes prefill and decoding consume the device sequentially. Real sales also depend on queueing, latency, caching, quality, uptime, distribution and demand.

04 / READ THE DATA

The conditions matter.

01

Prefill ≠ generation

Prefill reads input. Generation writes output. A fast result in one stage cannot stand in for the other.

02

Keep recipes together

Compare model, chip, quantization, context, cache state, prompt and concurrency. Different recipes appear as separate points.

03

Commercial use needs its own check

Flash-Next's model license requires separate permission for Model as a Service. Price snapshots and theoretical capacity are not contracted revenue.

MODEL LICENSE ↗
CAPACITY MODELoutput tok/s = 1 / (input:output ratio / PP + 1 / TG)

A conservative sequential-time approximation. Multi-user serving should use measured aggregate throughput with a latency target.