Read the prompt
Cold prompt processing. Higher is faster, under the same context and recipe.
A living record of local AI inference—prefill, token generation, the kernels behind each move, and what the numbers could mean for selling tokens.
Latest sourced event: 28 Sep 2026 · Public log through 29 Sep 2026
First documented dates, in order. Lines connect only comparable recipes. A drop is part of the story.
Cold prompt processing. Higher is faster, under the same context and recipe.
One stream and aggregate multi-stream results are labelled separately.
From the first verified day forward—model by model, patch by patch.
A transparent capacity scenario, not observed demand or a promise of profit.
Capacity assumes prefill and decoding consume the device sequentially. Real sales also depend on queueing, latency, caching, quality, uptime, distribution and demand.
Prefill reads input. Generation writes output. A fast result in one stage cannot stand in for the other.
Compare model, chip, quantization, context, cache state, prompt and concurrency. Different recipes appear as separate points.
Flash-Next's model license requires separate permission for Model as a Service. Price snapshots and theoretical capacity are not contracted revenue.
MODEL LICENSE ↗output tok/s = 1 / (input:output ratio / PP + 1 / TG)A conservative sequential-time approximation. Multi-user serving should use measured aggregate throughput with a latency target.