Paged Attention and Memory Efficiency in Production LLM Serving
How PagedAttention solved the memory crisis that made LLM serving uneconomical.

Paged Attention and Memory Efficiency in Production LLM Serving.
Why KV cache memory was the bottleneck that made high-concurrency LLM serving unworkable before PagedAttention
Every time a transformer generates a token, it needs to remember the key and value tensors from every token that came before it in that sequence, for every layer in the model. That memory, the KV cache, has to stay resident for the entire life of the request; there's no discarding it mid-generation and recomputing it later without paying a heavy penalty. The trouble is that this cache isn't a fixed size. It grows with every token the model produces, so a serving system has to decide up front how much room to set aside, and getting that decision wrong in either direction is expensive.
Before PagedAttention, most systems solved this by over-provisioning. The vLLM paper, presented at SOSP 2023, found that storing KV caches in contiguous memory wasted 60 to 80 percent of KV cache memory due to fragmentation and over-reservation KV Cache Optimization: Memory Efficiency for Production LLMs | Introl…. That's the difference between a GPU that can serve dozens of concurrent users and one that chokes on a handful.
The numbers get more concrete once you size a real model. Take a model with 80 layers and 8 KV heads at 128 dimensions each, storing both K and V in FP16: that works out to 327,680 bytes of KV cache per single token jarvislabs.ai. A single prompt at 4,000 tokens already costs 1.34 GB jarvislabs.ai. Pushing the context window to 8,000 tokens causes a single request to run roughly 20 GB KV Cache Optimization: Memory Efficiency for Production LLMs | Introl…. Now multiply that across a batch of 32 concurrent requests, a perfectly ordinary batch size for a production endpoint, and total demand reaches something like 640 GB, more memory than the model's own weights consume KV Cache Optimization: Memory Efficiency for Production LLMs | Introl….
This is why fragmentation wasn't a cosmetic inefficiency. It was the hard ceiling on batch size, and batch size is the single biggest lever a serving system has over throughput and cost per request. The consequence ran in two directions at once jarvislabs.ai. Low batch size meant poor GPU utilization, driving cost per request up. And when memory pressure did build, systems responded by preempting in-flight requests to free space, which showed up to users as latency spikes with no warning. High-concurrency serving wasn't merely inefficient under this regime; it was economically unworkable at any scale that mattered. Concrete scale of the problem (sizing a 70B-class model).
How PagedAttention borrows OS paging to eliminate KV cache fragmentation
PagedAttention was introduced in 2023 by Woosuk Kwon and collaborators, in the paper that also introduced the vLLM serving engine, presented at SOSP that year.
Each request keeps a logical block table, a lightweight map from logical positions in the sequence to the physical memory pages actually holding that data. The attention kernel resolves this mapping at runtime, on the fly, as it computes each step. Because blocks are fixed size, internal fragmentation drops to almost nothing; the only waste left over is the unused tail of the very last block in a sequence. And because allocation happens on demand as decoding proceeds rather than up front for a worst-case length, the system never reserves memory for tokens that might never get generated.
The sharing mechanism follows naturally from this design. When two requests share an identical prefix, say the same system prompt, both can point their block tables at the same physical blocks rather than each holding a private copy. A new block only gets allocated once a sequence diverges from that shared prefix, a scheme borrowed directly from copy-on-write memory management in operating systems. The original SOSP 2023 paper reports this brings KV cache waste down to near-zero, under 4 percent, while enabling this flexible sharing both within and across requests.
The OS analogy is not just a teaching device; it's the actual mechanism. A process running on a laptop doesn't need physically contiguous RAM to execute; the operating system's page table handles the indirection invisibly. What PagedAttention does not touch matters just as much. The actual math of attention, the dot products, the softmax, the weighted sum over values, comes out bit-for-bit identical to a naive implementation. A deeper cause produces this: the memory management layer sits underneath the computation, not as an approximation of it, and that layering is what makes the behavior possible. Nothing about model quality changes; only how efficiently the hardware holds onto the numbers involved. The central insight is to borrow virtual memory and paging from OS design, storing the KV cache in fixed-size blocks that map to non-contiguous physical GPU memory rather than requiring one large contiguous allocation. The throughput outcome was a 2–4× improvement over FasterTransformer and Orca at equivalent latency, with the improvement more pronounced with longer sequences, larger models, and complex decoding algorithms, per the SOSP 2023 paper.
How vLLM turned PagedAttention into a production serving engine
vLLM is the open-source engine built to put PagedAttention into practice, and a 2024 survey of LLM serving systems described the technique as having become an industry norm, with support appearing across TGI, vLLM itself, and TensorRT-LLM. That's a fairly rare trajectory for a research paper: from SOSP presentation to a load-bearing piece of infrastructure across competing serving stacks in roughly a year KV Cache Optimization: Memory Efficiency for Production LLMs | Introl….
The project has kept moving. Under the hood, the current release line runs on what's called the V1 engine, which became the default in 2025 and replaced the original engine core wholesale, its scheduler, KV cache manager, worker, sampler, and API server all rebuilt. In throughput terms, at equivalent configurations, vLLM has been measured at 793 tokens per second against 41 tokens per second for Ollama, a gap that reflects just how much continuous batching and paged memory management change the economics of serving vLLM Production Deployment | Introl Blog.
Production adoption backs this up. vLLM runs workloads at Meta, Mistral AI, Cohere, and IBM. On the performance frontier, vLLM has reached a 2026 milestone of 5,000 tokens of throughput and 180 interactivity running Qwen3.8-2.4T on GB300 NVL72 hardware with prefill-decode serving. And the ecosystem keeps extending sideways: vllm-metal ports vLLM's paged, continuously batched serving stack to Apple Silicon, with flatter time-to-first-token under concurrent agent load, batched speculative decoding, and automatic prefill acceleration on M5 chips.
This reads less like a research artifact that got cited a lot and more like the dominant open-source inference runtime, and the adoption pattern across hyperscalers, startups, and hardware vendors alike is the clearest validation available that the original PagedAttention premise was correct. Hardware breadth spans NVIDIA GPUs, AMD GPUs, Intel GPUs, x86/ARM/PowerPC CPUs, plus hardware plugins for Google TPUs, Intel Gaudi, IBM Spyre, Huawei Ascend, Rebellions NPU, Apple Silicon, MetaX GPU, and others. Stripe achieved a 73% inference cost reduction via vLLM migration, handling 50M daily API calls on one-third of its prior GPU fleet vLLM Production Deployment | Introl Blog KVServe: Service-Aware KV Cache Compression for Communication-Efficie…. In January 2026, inference startup Inferact raised $150M to commercialize vLLM vLLM Production Deployment | Introl Blog en.wikipedia.org KVServe: Service-Aware KV Cache Compression for Communication-Efficie….
The costs vAttention's design creates and vAttention's proposed alternative
Success doesn't mean the design is free of tradeoffs, and a 2025 paper on a system called vAttention, presented at ASPLOS, made the architectural case. Its argument: PagedAttention's non-contiguous block layout isn't something you bolt onto an existing attention kernel. It requires rewriting the kernel itself to understand paging, to walk the block table, to resolve indirection at every step of computation.
That requirement carries a real cost. Every new attention kernel variant, a FlashAttention update, a new hardware backend, has to be reimplemented with paging awareness baked in, which the vAttention authors frame as a tax on software complexity and portability. It's redundant engineering effort in a specific sense: paging logic that, in principle, belongs at the memory management layer ends up duplicated inside the compute kernel instead, and the block-table lookups themselves add execution overhead at runtime.
vAttention's proposed fix inverts the approach. Rather than paging the KV cache explicitly inside the kernel, keep it contiguous in virtual memory and let the operating system's own demand paging handle physical allocation underneath. That separates memory management cleanly from attention computation, so kernel authors never have to think about paging at all.
Whether this wins out in practice is an open question. PagedAttention has enormous ecosystem momentum behind it: tooling, hardware support, and a long list of production deployments that vAttention, as an architecture, doesn't yet have anywhere close to matching. What the disagreement reveals, regardless of which side eventually wins more adoption, is that the block-table indirection making PagedAttention so effective at the memory layer also couples memory management to compute in a way that becomes a genuine maintenance and portability burden as kernels keep evolving. For anyone actually operating inference infrastructure, the lesson is to recognize the tradeoffs at play rather than pick a side today: memory management architecture for LLM serving is still an open design problem, not a solved one, no matter how dominant vLLM has become in the meantime.
Prefix caching and KV sharing: how systems reuse computed cache across requests
PagedAttention's copy-on-write sharing opens the door to a broader technique known as prefix caching. If two requests share an identical prefix, a system prompt, a set of few-shot examples, a chunk of retrieved context in a RAG pipeline, the physical KV blocks covering that shared prefix get computed exactly once and reused, with no recomputation for the second request or the hundredth. On a warm cache in production, hit rates around 87 percent have been reported, and the metrics to watch to manage this in practice include cache usage percentage, prefix cache hit rate, eviction rate, and effective cache throughput backend.ai.
The catch is that conventional prefix caching only helps when the repeated content is at the front of the prompt. Real workloads don't cooperate with that assumption nearly as often as system designers might hope. Mid-conversation history repeats. Retrieved document chunks appear at different positions across turns. Shared reasoning artifacts get passed between agents in a multi-agent pipeline, none of it aligned neatly at a shared prefix boundary.
A system called SparseX, from MemTensor and dated 2026, targets exactly this gap. Rather than treating the prefix as the only reusable unit, it works over contiguous token segments wherever they occur in a sequence. It introduces what it calls Sparse-Q indices to flag which tokens need context-dependent correction once a segment gets reused out of its original position, then performs Sparse-KV recomputation within a single forward pass to restore the cross-segment interactions that pure caching would otherwise lose. Early layers keep full attention, for a stable importance signal, while later layers switch over to the sparse recomputation approach. The system is model-agnostic and requires no retraining, and it's built to sit on top of existing PagedAttention and prefix cache infrastructure rather than replace it, implemented concretely as SparseX-vLLM. Its natural targets are multi-round chat, retrieval-augmented generation, and multi-agent workflows, precisely the settings where repetition occurs away from the prompt's front edge.
Other work chips at adjacent corners of the same problem. HotPrefix, published in the Proceedings of the ACM on Management of Data in 2025, builds a hotness-aware scheduler that prioritizes which prefixes stay cached based on how often they get reused. DroidSpeak, also from 2025 and slated for NSDI 2026, tackles a different edge case: sharing KV cache across fine-tuned variants of the same base model, rather than assuming every request runs against an identical checkpoint. Taken together, these systems point at a simple operating principle: prefix hit rate is one of the highest-leverage numbers a serving team can watch, since a system running at 87 percent hit rate on a warm cache is effectively multiplying its available GPU compute budget for every request that falls in that 87 percent. vLLM's prefix caching enables a 400%+ utilization improvement for standardized prompts, while cross-instance KV cache sharing delivers a 3–10× latency reduction for repetitive workloads KVServe: Service-Aware KV Cache Compression for Communication-Efficie… backend.ai.
Eviction, compression, and the challenge of fitting more KV cache into less memory
Once a KV cache exists, a serving system has two basic ways to make more of it fit in fast memory jarvislabs.ai. It can evict blocks to slower storage, holding onto them for later reuse if the request comes back, or it can compress the blocks so each one takes less room in the first place. Historically, most systems picked one of these and stuck with it, treating eviction and compression as separate subsystems rather than parts of a single decision.
A paper called EVICPRESS, released in December 2025 out of a group spanning the University of Chicago, UC Berkeley, Tensormesh, MIT, UC Santa Cruz, Stanford, and Microsoft, argues that treating them separately leaves performance on the table. Its logic: a KV cache that can survive aggressive compression without losing generation quality should get compressed, while one that can't tolerate that loss should either stay in fast storage or get evicted losslessly instead. EVICPRESS builds a unified utility function to quantify this tradeoff, a profiling module that periodically scores every combination of eviction and compression configuration across all the concurrent contexts a system is juggling, then places blocks across storage tiers using a fast heuristic informed by those scores. Compression sensitivity isn't uniform across contexts, and a system that applies one compression policy to every KV cache indiscriminately is giving up quality and latency it didn't have to give up.
A cluster of other systems attacks adjacent slices of the same problem. PagedEviction, from 2025, is a block-wise eviction algorithm built specifically for PagedAttention's layout, identifying and dropping low-importance blocks without touching the CUDA kernel underneath. Entropy-guided caching allocates cache budget by measuring attention entropy layer by layer, giving more room to layers whose attention spreads broadly and less to layers with tightly focused attention patterns. SnapKV, also NeurIPS 2024, works by letting the model identify what it's actually going to need before generation even starts, trimming what gets cached accordingly.
A tension runs through all of this. Compression buys memory headroom at the risk of quality loss; eviction protects quality but costs latency when the evicted data has to come back. EVICPRESS's real contribution isn't either technique individually, it's making the choice between them adaptive and per-context rather than fixed and global. The result is up to 2.19× faster time-to-first-token (TTFT) at equivalent generation quality, evaluated across 12 datasets and 5 models EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM…. KVQuant targets 10-million context length LLM inference with KV cache quantization and was published at NeurIPS 2024. Streaming LLM for infinite generation maintains initial "attention sink" tokens (the first 4–8 tokens) plus a sliding window of recent tokens, enabling fixed-memory generation of theoretically unlimited length, though quality degrades for tasks requiring long-range dependencies.
KV cache compression under disaggregated serving, where the network is the bottleneck
Disaggregated serving splits prefill and decode across separate GPU nodes, or offloads KV cache entirely to a remote pool, and the moment that split happens, KV cache becomes an explicit payload that has to cross network and storage boundaries rather than remaining a purely local memory management concern. That changes the nature of the bottleneck entirely.
The scale of the bandwidth problem is stark. Llama 3.1-70B generates 39.06 GB of KV cache at a 128,000-token context spheron.network. Serving 32,000-token requests with Qwen3-235B across a 64-node prefill cluster requires 2.1 terabits per second of KV egress bandwidth, at a time when common cross-cluster cloud links top out below 100 gigabits per second, and remote storage or KV pool throughput often runs below 10 gigabits per second spheron.network. Under those constraints, KV communication time can eat up to 60 percent of total job completion time in a PD-separated deployment, according to measurements behind the KVServe system.
Its framing: KV compression under disaggregation is a decision that shifts as workload conditions shift. It unifies existing compression methods into a modular strategy space that supports mixing techniques across methods rather than locking into one. At runtime, a service-aware online controller pairs an analytical latency model with a lightweight bandit algorithm, choosing among the profiled candidates based on current service-level and bandwidth constraints, correcting for whatever gap exists between the offline profile and the messier online reality.
The observation that produces every single multiplier matters more than any one of them on its own: the latency-optimal compression choice moves depending on workload and bandwidth regime, so a static configuration tuned for one operating point can actively increase latency at another. Anyone building disaggregated inference infrastructure has to treat compression as a live, adaptive decision. KVServe, arXiv:2605.13734, May 2026, is a collaboration among the University of Chinese Academy of Sciences, the Institute of Computing Technology CAS, and Shanghai Jiao Tong University. The Bayesian Profiling Engine searches the strategy space and distills a 3D Pareto candidate set, reducing offline search overhead by 50×. The results show up to a 9.13× job completion time speedup in PD-separated serving and up to a 32.8× TTFT reduction in KV-disaggregated serving.
Sources
- EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
- SparseX: Efficient Segment-Level KV Cache Sharing for Interleaved LLM Serving
- KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving
- Efficient Memory Management for Large Language Model Serving with PagedAttention | Proceedings of the 29th Symposium on Operating Systems Principles
- VLLM
- KV Cache Optimization: Memory Efficiency for Production LLMs | Introl Blog
- vLLM Production Deployment | Introl Blog
- [2405.04437] vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention


