Posts
Archive
Tags

Inference

KV cache, and why LLM inference is memory-bound

The cache that makes autoregressive decoding fast also makes it the thing that runs out of memory first.

February 8, 2026 · 2 min · mc

© 2026 mc · notes · Powered by Hugo & PaperMod