Season one · stop 5 · The KV cache

Why It Doesn't Reread Your Chat

“Every new word looks back at everything before it. So why doesn't a long chat grind to a halt? Because the model keeps notes. We timed it with the notes and without them.”

Unpacking the world: 1.8 MB of drawings and measured data.

AgenticAmit A field guide to the KV cache

Field guide · the KV cache

Why it doesn't reread your whole chat.

A model writes one piece at a time, and every piece looks back at everything before it. That should get slower with every word. It mostly doesn't, because the model keeps notes. We timed GPT-2 with the notes and without them.

Scroll to open the drawer ↓

The desk · one replyGPT-2 small · temperature 0.7
One piece at a time.
One index card
Every piece files cards.
The piece writes:
key

What this card is about. Newer pieces match their question against it.

value

What this card says, once it's picked.

queryThe newest piece brings a question of its own, and checks it against every key in the drawer.
Without notesfirst 24 steps
Reread everything, every time.
piece run through the model
0pieces run through the model so far
With notes · the KV cachefirst 24 steps
Write one card, read the rest.
new cards writtencards read from the drawer
0pieces run through the model so far
Pinned · why old cards never changeNIPS 2017
Vaswani et al.: • Similarly, self-attention layers in the decoder allow each position in the decoder to attend to all positions in the decoder up to and including that position. We need to prevent leftward information flow in the decoder to preserve the auto-regressive property. We implement this inside of scaled dot-product attention by masking out (setting to −∞) all values in the input of the softmax which correspond to illegal connections. See Figure 2.

Vaswani et al. · Attention Is All You Need

Proof 1 · same reply?1,000 pieces each way
The notes change nothing.
Proof 2 · time for each new piece
Without notes, every word costs more.

The catch · a current open model
The drawer grows with every word.

Pinned · the published configHugging Face
Qwen2.5-7B-Instruct config.json: "num_attention_heads": 28, "num_hidden_layers": 28, "num_key_value_heads": 4,config.json: "hidden_size": 3584,Qwen2.5-7B-Instruct model card: Number of Parameters: 7.61Bmodel card: Context Length: Full 131,072 tokens and generation 8192 tokens Please refer to this section for detailed instructions on how to deploy Qwen2.5 for handling long texts.

huggingface.co/Qwen/Qwen2.5-7B-Instruct · retrieved 28 Sep 2026

Pinned · the same sum, 2023SOSP 2023
Kwon et al.: Large KV cache. The KV Cache size grows quickly with the number of requests. As an example, for the 13B parameter OPT model [62], the KV cache of a single token demands 800 KB of space, calculated as 2 (key and value vectors) × 5120 (hidden state size) × 40 (number of layers) × 2 (bytes per FP16). Since OPT can generate sequences up to 2048 tokens, the memory required to store the KV cache of one request can be as much as 1.6 GB. Concurrent GPUs have memory

Kwon et al. · PagedAttention (vLLM)

Fix 1 · grouped-query attention
Share the drawers.
Pinned · the ideaEMNLP 2023
Ainslie et al. Figure 2 (diagram): Multi-head Grouped-query Multi-query Values Keys QueriesAinslie et al.: Figure 2: Overview of grouped-query method. Multi-head attention has H query, key, and value heads. Multi-query attention shares single key and value heads across all query heads. Grouped-query attention instead shares single key and value heads for each group of query heads, interpolating between multi-head and multi-query attention.

Ainslie et al. · GQA

Fix 2 · PagedAttentionvLLM
Don't reserve the whole drawer.
Pinned · the waste, measuredSOSP 2023
Kwon et al.: allocated size can be different for each request. Indeed, our profiling results in Fig. 2 show that only 20.4% - 38.2% of the KV cache memory is used to store the actual token states in the existing systems.Kwon et al.: ory usage. Our evaluations show that vLLM improves the throughput of popular LLMs by 2-4× with the same level of latency compared to the state-of-the-art systems, such as FasterTransformer and Orca. The improvement is more

Kwon et al. · PagedAttention (vLLM)

Keeping the drawer · between messagesmeasured
The slow part is filling it.
Pinned · a discount for reused notesAnthropic docs
Anthropic prompt caching docs:  The previous table reflects the following pricing multipliers for prompt caching: 5-minute cache write tokens are 1.25 times the base input tokens price 1-hour cache write tokens are 2 times the base input tokens price Cache read tokens are 0.1 times the base input tokens price (see the table footnote for per-model exceptions) These multipliers stack with other pricing modifiers such as the Batch API discount and data residency. See pricing for full details.

platform.claude.com · prompt caching · retrieved 28 Sep 2026

Pinned · the same ideaOpenAI docs
OpenAI prompt caching guide: The prompt cache stores key-value (KV) tensors, not the tokens themselves.OpenAI prompt caching guide: Prompt caching reuses work when requests share the same prompt prefix. This provides three main benefits: Compute-efficient: Avoid recalculating a prompt prefix that the model has already processed. Cheaper input tokens: Pay the model’s reduced cached-input rate for reused tokens, discounted up to 90%. Faster: Reduce the time spent processing input before the response starts.

developers.openai.com · prompt caching · retrieved 28 Sep 2026

Take this with you

The model keeps notes. The notes are why long chats cost memory.

Amit, calm and reassuring