Anatomy of an inference request
What happens between hitting enter and seeing the first token, and why the two numbers that describe it are set by completely different machinery.
Every request to a language model API goes through the same four stages: tokenize, prefill, decode, stream. Most conversations about "model speed" go wrong because they treat these as one thing. They are not. The time you wait for the first token and the rate at which the rest arrive are produced by different parts of the system, they fail differently, and they cost differently.
Stage one: tokenize
Your prompt arrives as text and the model cannot read text. It reads tokens, integer IDs from a fixed vocabulary that was frozen when the model was trained. A tokenizer chops your string into those IDs. This step is cheap and fast, microseconds against everything that follows, but it decides the units everything downstream is measured in. Context limits, pricing, and latency are all counted in tokens, not characters or words.
"The KV cache is stored in GPU memory"
|
v
[The] [ KV] [ cache] [ is] [ stored] [ in] [ GPU] [ memory]
|
v
one integer ID per piece, from a fixed vocabulary
The same text always produces the same IDs. Paste a prompt into any public tokenizer playground and the segmentation you see is exactly what the model receives.
Stage two: prefill
Now the model reads your prompt. All of it, in one pass. The serving system runs every prompt token through the network at once, and because each token's computation is independent at this stage, the work parallelizes beautifully across a GPU. Prefill is compute-bound: the accelerator is genuinely busy doing matrix multiplication.
Prefill produces two things. The first is the probability distribution for the very first output token. The second matters more for cost: the model saves the intermediate attention state (the keys and values) for every prompt token, so it never has to re-read the prompt again. That saved state is the KV cache.
The KV cache is why long prompts cost what they cost. It is a block of GPU memory, sized to your context length, that must stay resident for the entire life of your request. A conversation with a hundred thousand tokens of history is holding a large, expensive allocation on hardware measured in tens of gigabytes, whether the model is generating or just waiting for your next message in a multi-turn exchange.
Time to first token, TTFT, is the cost of prefill plus queueing. Long prompt, longer prefill, later first token. When your app feels slow to start responding, this is the stage to look at.
Stage three: decode
Generation is nothing like prefill. The model produces one token at a time, and each new token depends on everything before it, so there is no parallelizing across the sequence. Each step is one full forward pass that reads the entire KV cache and appends one entry to it. The GPU spends most of a decode step moving weights and cache through memory rather than doing arithmetic. Decode is memory-bandwidth-bound.
At the end of each step the model has not chosen a token. It has produced a probability distribution over the whole vocabulary, and a sampler picks from it.
prompt so far: "The KV cache is stored in GPU"
memory ████████████████████████ 72%
RAM ████ 11%
hardware ██ 6%
memory, █ 4%
VRAM █ 3%
(illustrative, not measured)
Tokens per second, the throughput number, is a property of decode. It depends on model size, hardware, batch load, and cache length. Notice that nothing about it improves your TTFT, and nothing about prefill speed improves your streaming rate. Two numbers, two subsystems.
Other people's traffic
Everything so far assumed your request had the GPU to itself. In production it never does. Schedulers like the one vLLM popularized use continuous batching: at every decode step, the server assembles a batch from all active requests, generates one token for each, then reassembles the batch. Requests join and leave the batch mid-flight, which is what keeps expensive GPUs busy.
The consequence for you: your decode speed varies with load you cannot see. The same request, same prompt, same model can stream at noticeably different rates at different hours, because the batch you share changes. If you are benchmarking a provider with five requests on a Tuesday afternoon, you are measuring the weather as much as the system.
Why streaming exists. Tokens become available one at a time as decode steps complete, and sending each one immediately is the natural shape of the system. Waiting for the full response and delivering it at once is the artificial version.
What this means when you build
A few practical readings fall straight out of the anatomy.
Perceived speed is mostly TTFT, and TTFT is mostly your prompt. Trimming a bloated system prompt does more for how fast your product feels than switching to a model with a higher tokens-per-second headline.
Input and output tokens are priced differently because they are different work. Prefill parallelizes; decode is a serial grind that occupies cache memory the whole way through. The price sheet reflects the physics.
Prompt caching is KV cache reuse. When providers offer cached input pricing, they are letting a stored prefix of the KV cache stand in for redoing prefill. This only works if your prompt's opening bytes are identical across requests, which is why the ordering of your prompt is an engineering decision. Stable content first, variable content last.
Long conversations get slower and pricier for structural reasons. Every turn re-reads a longer history during prefill (or holds a bigger cache), and every decode step attends over more entries. Nothing is wrong with your integration. The machine is doing more work.
None of this requires touching a GPU to reason about. The published architecture is enough, and once the four stages are distinct in your head, most provider docs, pricing pages, and latency dashboards stop being mysterious.