How Do KV Cache and Paged Attention Speed LLMs?

192.2K views
•
June 30, 2026
by
IBM Technology
YouTube video player
How Do KV Cache and Paged Attention Speed LLMs?

TL;DR

KV cache speeds autoregressive generation by storing previously computed key and value matrices, allowing each new token to reuse its attention history instead of recomputing it. Paged attention improves serving capacity by allocating that cache in small, noncontiguous blocks on demand, reducing fragmentation and wasted GPU memory across concurrent requests.

Transcript

You've got a large language model running. With one user, time to first token is super quick. And at 10 users, we're starting to see latency climb. But then at a hundred, you're watching GPU memory spike and throughput tank. And every cycle is wasted money. Your model is most likely not the culprit here, but rather how your memory is being used dur... Read More

Key Insights

  • LLM inference consists of a compute-bound prefill phase and a memory-bound decode phase. Prefill processes the complete prompt through every transformer layer, while decode generates tokens individually and repeatedly retrieves the growing context stored in GPU memory.
  • KV cache is a memory-for-compute trade-off that stores key and value matrices from previous generation steps. Each new token computes only its own query, key, and value, then attends over the cached history instead of recalculating earlier keys and values.
  • GPU memory is the central bottleneck in large-scale LLM serving. The example 13-billion-parameter model requires about 26 GB for weights on a 40 GB A100, consuming 65% of its memory before any request reaches the model.
  • Traditional cache allocation wastes memory by reserving a fixed contiguous region based on the maximum output length. With a 2,048-token limit and a typical 200-token input plus 300-token output, memory for 1,500 reserved tokens remains unused for that request.
  • Internal fragmentation is unused capacity inside an oversized allocation, while external fragmentation consists of gaps between allocations created by requests with different lengths. A system can therefore have enough total free memory for a request but lack one sufficiently large contiguous region.
  • Paged attention stores KV cache data in small fixed pages containing 16 tokens by default. Pages are allocated as needed and may reside anywhere in GPU memory, while a lightweight block table maps the logical addresses seen by the model to physical VRAM locations.
  • Prefix caching reduces duplicate work when requests share token sequences such as a system prompt. Paged attention hashes each KV block, allowing multiple requests to reference the same physical memory and skip shared prefill work that has already been computed and stored.
  • Chunked prefill protects ongoing decode work from long prompts by serving decode requests first and using the remaining compute budget for prompt chunks. The transcript reports up to 50% throughput improvement in production and suggests setting maximum batch tokens above 2,048 alongside it.

Explore YouTube Video Summarizer or Get YouTube Transcript Extractor

Questions & Answers

Q: How does KV cache speed up LLM token generation?

KV cache stores the key and value matrices calculated for tokens from earlier generation steps. Without it, producing a later token requires recalculating keys and values for every preceding token. With the cache, each new token computes only its own query, key, and value, then attends over the stored history. This exchanges additional GPU memory usage for less repeated computation.

Q: What is the difference between prefill and decode in LLM inference?

Prefill occurs before the first output token and processes the complete prompt through every transformer layer to construct its mathematical representation. It is described as compute bound. Decode follows prefill and produces output one token at a time. It repeatedly retrieves the request's growing context from GPU memory, so it is memory bound and depends heavily on efficient KV cache management.

Q: Why does LLM latency increase when concurrent traffic grows?

Every concurrent request maintains a growing context that must be stored and retrieved during token generation. As request volume rises, GPU memory can fill or become fragmented, and repeated or inefficient context handling creates visible latency. Traditional fixed allocations also reserve capacity that many requests never use, reducing the number of active requests the GPU can support and harming overall throughput.

Q: How does paged attention reduce GPU memory waste?

Paged attention divides each request's KV cache into small fixed-size pages instead of reserving one large contiguous block. The default page holds 16 tokens, pages are allocated only when needed, and they can occupy noncontiguous physical locations in VRAM. A lightweight block table maps logical pages to physical pages, reducing both unused reserved space and allocation gaps.

Q: What are internal and external KV cache fragmentation?

Internal fragmentation occurs when a request receives an allocation based on the maximum sequence length but uses only a small part of it. External fragmentation occurs when requests of different lengths leave gaps between allocations. Even when the combined free memory is large enough for a new request, allocation can fail if no single contiguous region can satisfy its requirement.

Q: How should GPU memory utilization be configured in vLLM?

GPU memory utilization controls the fraction of available VRAM assigned to KV cache, and the stated default is 0.9. The transcript recommends considering 0.95 for stable workloads to accommodate more concurrent requests, or reducing it to 0.8 when burst traffic causes out-of-memory errors. The specific model and workload should be benchmarked before committing to a setting.

Q: When does prefix caching improve LLM inference?

Prefix caching helps when requests repeatedly share the same token sequence, such as a system prompt. Paged attention hashes each KV block so those requests can point to one physical copy rather than storing duplicates. The transcript says cache hit rates of 75–95% are common in RAG pipelines, multi-turn chat, and coding agents, allowing shared prefill work to be skipped.

Q: When should speculative decoding be used?

Speculative decoding is most useful when interactive latency matters more than maximum throughput. During memory-bound decode, a smaller draft model proposes several tokens and the larger model verifies them in one forward pass, correcting rejected proposals while preserving mathematically identical output quality. Its gains become smaller at very high concurrency because batching already keeps the GPU busy.

Summary & Key Takeaways

  • LLM inference has two phases with different bottlenecks. Prefill processes the entire input through every transformer layer before producing the first token, making it compute bound. Decode generates tokens one at a time while repeatedly retrieving the growing context from GPU memory, making efficient memory access critical to latency and throughput.

  • KV cache avoids repeatedly calculating keys and values for earlier tokens during autoregressive generation. However, conventional serving reserves a contiguous allocation based on the maximum sequence length for every request. Shorter requests leave much of that reservation unused, while variable request lengths create gaps that can prevent otherwise sufficient memory from being allocated.

  • Paged attention divides KV cache storage into fixed-size pages, which contain 16 tokens by default and can occupy noncontiguous locations in GPU memory. A block table connects logical pages to their physical VRAM locations. Deployments can additionally tune memory utilization, cache shared prefixes, use chunked prefill, and enable speculative decoding.


Read in Other Languages (beta)

Share This Summary 📚

Explore More Summaries from IBM Technology 📚