Why it doesn't reread your whole chat.
A model writes one piece at a time, and every piece looks back at everything before it. That should get slower with every word. It mostly doesn't, because the model keeps notes. We timed GPT-2 with the notes and without them.
Scroll to open the drawer ↓
One piece at a time.
Every piece files cards.
What this card is about. Newer pieces match their question against it.
What this card says, once it's picked.
Reread everything, every time.
Write one card, read the rest.
Vaswani et al. · Attention Is All You Need
The notes change nothing.
Without notes, every word costs more.
The drawer grows with every word.



huggingface.co/Qwen/Qwen2.5-7B-Instruct · retrieved 28 Sep 2026
Kwon et al. · PagedAttention (vLLM)
Share the drawers.

Ainslie et al. · GQA
Don't reserve the whole drawer.

Kwon et al. · PagedAttention (vLLM)
The slow part is filling it.
platform.claude.com · prompt caching · retrieved 28 Sep 2026

developers.openai.com · prompt caching · retrieved 28 Sep 2026
The model keeps notes. The notes are why long chats cost memory.