Storage Scaling Bottlenecks in Large LLM Inference Clusters
KV cache grows without bounds, forcing engineers to choose which storage tier fails first.

Every scaling problem in a large language model cluster eventually becomes a storage problem, and the KV cache is why. When a model generates text one token at a time, it must attend to every token that came before it at each step of decoding. Rather than recompute the key and value vectors for that entire history on every step, inference engines store them in a structure called the KV cache, trading GPU memory for GPU compute. That trade is efficient, but it is not free, and it does not stay small. The size of the KV cache grows linearly with context length, batch size, model width, and the number of layers in the model, so any increase in context window or concurrent users multiplies the memory demand directly. Given that trajectory, if you are scaling an inference cluster, the real engineering question is which tier in that hierarchy gives out first, and under what conditions, because each tier fails differently and the fix for one failure often creates the next one.
How GPU HBM becomes the first tier to saturate
High-bandwidth memory on the GPU is where this pressure lands first, and it lands hardest, because HBM is both the fastest tier available to the model and the smallest relative to what long-context and multi-turn workloads demand of it. During decoding, the model attends to the full KV cache at every single generation step, and that operation is bound by memory bandwidth rather than by raw compute throughput. In practice, this means HBM bandwidth, not FLOPS, sets the ceiling on how fast tokens get generated. As context length grows, the KV blocks for active requests start competing directly with the model's own weights for the same pool of HBM capacity, and once that pool fills, the serving system has no choice but to evict or offload cache blocks elsewhere, which is what pushes the problem down into the slower tiers of the hierarchy. Engineering teams have real tools to push back against this pressure: quantizing the KV cache from BF16 down to FP8 cuts bytes per token, paged memory allocation through PagedAttention reduces the fragmentation that wastes capacity, and newer attention variants like MLA shrink the cache footprint directly. None of these change the underlying linear scaling relationship between context and memory use. They buy time before HBM saturates, and they do not stop it from saturating eventually. Once HBM fills, the system has to go somewhere else, and that somewhere else is CPU DRAM.
Offloading KV Cache to CPU DRAM: Relocating the Problem to PCIe
Moving the KV cache into CPU DRAM looks like an obvious fix, and in terms of raw capacity it is: DRAM pools give a cluster orders of magnitude more room than HBM alone. This move turns the memory-capacity problem into something else rather than making it disappear. It turns it into a bandwidth problem at the PCIe interconnect that links the GPU to the host, and PCIe bandwidth is orders of magnitude narrower than HBM bandwidth. The paper "Understanding Bottlenecks for Efficiently Serving LLM Inference With KV Offloading" gives engineers a precise way to reason about this: it defines a critical ratio, κ_crit, as the cached-to-prefill token threshold at which execution flips from compute-bound to memory-bound. The measured consequence is stark: end-to-end latency in these offloaded setups is dominated by PCIe transfer time, accounting for 99% of it, while GPU compute utilization drops to roughly 28% of rated TDP. That is close to a full inversion of the assumption most cluster capacity planning is built on, which treats GPU compute as the scarce resource and the interconnect as incidental. Iteration-level schedulers are good at filling idle compute cycles with other work, but there is no other work to give the GPU when it is simply stalled waiting on a transfer to finish. Hardware changes such as NVLink C2C links or unified HBM architectures do raise κ_crit meaningfully, so a server gets more headroom before it tips into the memory-bound regime, but these are physical changes to server architecture, not something a software team can configure its way into, and only a small minority of deployed clusters have them today.
KV Cache on NVMe: Mechanics and Tradeoffs
NVMe SSDs are the next tier down from DRAM, and their appeal is straightforward: cost per gigabyte low enough to make very large KV working sets economically realistic in a way that DRAM alone cannot match. But whether an NVMe cache hit actually helps a given request depends on more than the device's rated bandwidth. It depends on the granularity at which data gets transferred, on when in the request's lifecycle that transfer gets scheduled, and on how the length of a cached prefix compares to the cost of simply recomputing it. Research behind py-kvcache characterized exactly this tradeoff inside vLLM, testing GPU, CPU, and NVMe tiers together against synthetic workloads, long-context benchmarks, and production traces, and the finding was that cache performance tracks transfer granularity, intermediate memory use, and scheduling timing far more closely than it tracks raw device bandwidth. That produces a genuine break-even point: for short prefixes, or on fast GPUs like the H100, the average request is actually cheaper to recompute from scratch than to fetch from disk. Where it does pay off is at long context lengths, where the py-kvcache implementation, using asynchronous direct I/O, bounded shared staging, and scheduler-aware preloading, loads from disk faster than LMCache; it also pays off on weaker GPUs, and in workloads with high prefix reuse across requests, such as retrieval-augmented generation, shared system prompts, or document question-answering. Results on LongBench and SCBench confirm the gains extend to irregular prefix chains and multi-turn conversations, but a production trace replay from Bailian also showed the opposite case: on an H100, GPU memory alone was enough to retain the prefixes that mattered, and NVMe offload added nothing. NVMe caching is a tool that earns its keep under specific conditions and costs something under others, and the job of an infrastructure team is to know which side of that line their own workload sits on.
Object storage as a KV tier for S3-compatible systems at runtime
Beyond NVMe, the outermost tier available to a cluster is networked object storage, and S3-compatible systems can extend KV cache capacity without any practical ceiling. The standard object storage abstraction does not match how an inference engine actually wants to consume KV data. S3 is built around object-level operations: name an object, fetch a byte range. A prefix cache hit, by contrast, returns many separate hash-addressed KV chunks, and the engine needs to consume them in strict layer order. If you issue a separate S3 request per chunk, fixed per-request overhead sits directly on the critical path of generation; if you fetch coarser objects instead, the system loses the fine-grained prefix reuse that made object storage worth reaching for. ObjectCache, described in the paper "Layerwise Object-Storage Retrieval for KV Cache Reuse," solves this by co-designing the storage protocol and the request scheduler together rather than treating them separately. KV cache is still stored as fine-grained, hash-addressed chunks, but the S3-compatible request itself is extended with a descriptor naming the matched chunks, the model's layout, and the order delivery needs to happen in, so that the storage server assembles and ships each layer directly to the serving node over RDMA. Long contexts dominate current production traffic, and for them this adds only modest latency compared to pulling the same data from local DRAM. That gap widens at shorter contexts, where there is less compute available on the GPU side to mask the transfer time. Object storage earns its place specifically when contexts run long enough that the savings from skipping prefill outweigh the cost of fetching the cache over the network. ObjectCache's scheduler also exploits the layer-by-layer transfer structure to overlap data delivery with compute across multiple concurrent requests, which keeps added time-to-first-token in check even when bandwidth is shared and capped across tenants. The broader signal here is that as context scale keeps growing and KV working sets outgrow what remote DRAM pools can hold, object storage becomes a credible architectural tier in its own right, but only when the infrastructure is built to expose its bandwidth over RDMA rather than through a generic HTTP interface.
Lossless vs. lossy compression tradeoffs across storage tiers
Compression cuts across every tier discussed so far, and whether it is lossless or lossy changes the economics of each one differently. This is not purely a question about model accuracy; it determines how much real capacity and bandwidth relief a given technique delivers, and which workloads can afford the trade involved in getting it. Lossy methods, quantizing the cache from BF16 down to FP8 or selectively evicting tokens through methods like H2O, SnapKV, or Ada-KV, cut the number of bytes stored per token substantially. That directly eases pressure on HBM capacity and reduces the volume of data that has to cross PCIe to DRAM or travel further out to NVMe and object storage, and those gains are concrete and measurable. For those cases, lossy compression introduces a risk to accuracy that is hard to bound in advance, which is the reasoning behind SplitZip, a lossless compression approach built specifically for KV cache transfer in prefill-decode disaggregated deployments. For storage tiers specifically, lossy compression shifts the break-even math for NVMe and object storage in a favorable direction, since fewer bytes moved means more requests clear the threshold where fetching beats recomputing, while lossless compression keeps that math unchanged but avoids introducing any precision risk. Infrastructure teams need to know which of those two regimes their workload actually falls into before picking a compression strategy, because if they choose wrong, they either leave bandwidth savings on the table or trade away correctness a workload cannot afford to lose.
High-Bandwidth Flash and CXL as emerging tiers that reorder the hierarchy
Two hardware approaches now under active research would each redraw this tier hierarchy, though they attack different parts of it and carry different risks if you deploy them carelessly. High-Bandwidth Flash, or HBF, stacks NAND flash behind a wide interface placed directly on the GPU package, promising capacity at flash scale with much lower read latency and higher bandwidth than an SSD reached over PCIe. If flash sits close enough to the GPU, the PCIe bottleneck that dominates today's KV offload latency goes away. The paper identifies three conditions a faster far tier needs to satisfy before it pays off: read I/O has to be the actual bottleneck, reads have to outweigh writes, and the bandwidth the device delivers has to be sustainable rather than transient. KV cache workloads in production violate these conditions in a specific way: writes outnumber reads, HBF's thermal limits are reached well below its peak rated bandwidth under a sustained write stream, and its TLC flash wears out faster than the SSD pool it was meant to replace, since it lacks the wear-leveling logic built into a dedicated SSD controller. HBF does earn a place when deployed selectively: directing long-lived, reusable KV data to HBF rather than transient cache data, budgeting writes deliberately, and coordinating around thermal limits lets a cluster capture HBF's bandwidth advantage without the write-amplification penalty that comes from treating it like a drop-in SSD replacement. A companion simulator, HBFSim, was released alongside this research so teams can model HBF-class device behavior under real GPU execution patterns before they commit to a deployment. CXL memory takes a different position in the hierarchy entirely: it is byte-addressable capacity inserted above NVMe, with latency far below flash. The ITME, or Inference Tiered Memory Expansion, research direction formalizes this as a way to expand memory capacity to the terabyte scale using a CXL-hybrid design, where a DRAM cache is backed internally by NVMe SSDs and a hardware-level prefetcher masks flash I/O latency rather than eliminating it. CXL and HBF are not competing solutions to the same problem. CXL extends the near tier, closer to DRAM, while HBF, deployed correctly, extends the far tier with bandwidth well above what NVMe offers. Both require workload-aware placement logic layered on top of the hardware to actually deliver the gains they promise on paper.
How disaggregated prefill-decode architectures change which storage tier is under pressure
Serving architectures that split prefill and decode across separate GPU nodes add a further wrinkle to all of this, because they change KV cache from something managed locally on one machine into something that has to move across a network between two different roles. In a disaggregated setup, the prefill node computes the initial KV cache for a prompt and then has to hand that cache off to a separate decode node that will carry the generation forward, so the transfer itself becomes part of the critical path for every single request rather than an occasional background operation. That reframes which tier in the hierarchy actually matters most for a given deployment. A cluster that disaggregates those two roles instead lives or dies by the network fabric connecting prefill and decode nodes, by how quickly the KV cache can be serialized, transferred, and reassembled on the decode side, and by whether compression techniques like SplitZip's lossless approach can shrink that transfer enough to keep it off the critical path. The storage hierarchy does not go away under disaggregation. It relocates to wherever the network boundary between roles sits, and the engineering discipline required to manage it, tracking which tier saturates first, matching compression strategy to workload, and designing scheduling that overlaps transfer with compute, applies just as much on that network boundary as it does inside a single node's memory hierarchy.
Sources
- Understanding Bottlenecks for Efficiently Serving LLM Inference With KV Offloading
- Building py-kvcache: A Performance Characterization of External KV Caching for vLLM with NVMe SSDs
- ObjectCache: Layerwise Object-Storage Retrieval for KV Cache Reuse
- HBF Sucks? A Full-Stack Characterization of High-Bandwidth Flash for KV-Centric LLM Serving
- SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving
- SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving


