Prefill ≠ generation
Prefill reads input. Generation writes output. A fast result in one stage cannot stand in for the other.
A living record of local AI inference—prefill, token generation, the kernels behind each move, and what the numbers could mean for selling tokens.
Latest sourced event: 28 Sep 2026 · Public log through 29 Sep 2026
The time axis runs earliest to latest. Every point still shows its PP, TG and test conditions.
Measured source-Mac throughput and a separate, unverified M5 Ultra scaling range. Source context stays attached to every value.
M5 Pro 16c → Ultra 80c: PP ×2–4, TG ×1.5–2.5. M5 Max 40c → Ultra 80c: PP ×1.2–1.8, TG ×1.2–1.7. Apple lists 307 / 614 / 1,200 GB/s memory bandwidth and 16 / 40 / 80 GPU cores for these specific variants. These broad factors allow for less than linear software scaling; they are scenarios, not cross-device benchmark findings. Each projection preserves its own context and engine.
Apple Pro / Max specs ↗ · Apple Ultra specs ↗| Model / metric | Short | 8K bucket | 16K bucket | 32K bucket |
|---|---|---|---|---|
| 35B PP · uncached | — | 2,575 | 2,373 | 2,011 |
| 35B TG · 1 stream | 210 | 156 | 149 | 143 |
| 27B PP · uncached | — | 398 | 395 | 363 |
| 27B TG · 1 stream | 74 | 55 | 55 | 54 |
SPEED-Bench coding prompts, reasoning on (27B medium), DFlash 2, HTTP, output cap 1,024; actual output lengths and TG cache state unpublished. Four short-stream aggregate: 35B 357; 27B 170. Cached 32K replay TTFT is a separate test. Inco official results ↗
| Metric / context | v13 | v14 | v15 | v16 | v17 |
|---|---|---|---|---|---|
| PP · 2K prompt | 4,014 | 4,097 | 4,196 | 4,224 | 4,180–4,300 |
| TG · 10 fresh, starting context unknown | 51.5 | 102.5 | 118.9 | 115.9 | 165.7 |
| TG · 2K coding, 1 stream | 62.0 | 75.0 | 124.2* | 85.9 | 179.4 |
| TG · 2 streams aggregate | — | — | — | 153.2 | 138.2 |
| TG · 4 streams aggregate | — | — | — | 220.3 | 159.1 |
All versions appeared in one post on 24 Sep (Taipei). Author confirmed MTP and said v17 batch is broken. Exact quantization, implementation, fresh-prompt context, output length, cache/reasoning and batch context were not published in the table. *v15 has an unexplained source footnote. Original table ↗ · Author reply ↗
Newest sourced results first; the oldest records are at the bottom. Each model and patch links to its source.
A transparent capacity scenario, not observed demand or a promise of profit.
Capacity assumes prefill and decoding consume the device sequentially. Real sales also depend on queueing, latency, caching, quality, uptime, distribution and demand.
Prefill reads input. Generation writes output. A fast result in one stage cannot stand in for the other.
Compare model, chip, quantization, context, cache state, prompt and concurrency. Different recipes appear as separate points.
Flash-Next's model license requires separate permission for Model as a Service. Price snapshots and theoretical capacity are not contracted revenue.
MODEL LICENSE ↗output tok/s = 1 / (input:output ratio / PP + 1 / TG)A conservative sequential-time approximation. Multi-user serving should use measured aggregate throughput with a latency target.