Tiered Memory Architecture for Long-Context LLM Inference
A tiered memory hierarchy solves long-context LLM serving bottlenecks.

Serving a long-context language model is, above everything else, a memory problem. The key-value cache that stores attention state grows linearly with both context length and batch size, and it does so fast enough to outgrow the model weights it was built to support. Serving a large model at very long contexts can require KV cache memory that simply will not fit on a single 80 GB GPU, no matter how the rest of the system is tuned. That's the arithmetic problem in its bluntest form: the cache is what runs out of room first.
Newer models make this worse by design. Context windows keep stretching further, and every added token multiplies the pressure on cache memory rather than adding to it at a flat rate. In ordinary production inference today, the KV cache already outweighs the model weights in memory footprint, which inverts the assumption most serving infrastructure was built around. Buying more GPU memory doesn't fix this, either, since HBM capacity is fixed per accelerator and scales only by adding entire GPUs, an expensive and often wasteful move. Host DRAM has its own ceiling too, set by how many sockets and memory channels a server actually has. And even the memory that gets allocated is often squandered: naive systems lose the majority of their assigned KV cache to fragmentation and over-allocation before a single extra token of context gets served.
None of this would matter as much if moving data between tiers were free. That gap remains unresolved, since moving data between tiers is not free. PCIe bandwidth is typically an order of magnitude lower than HBM bandwidth, so cross-tier transfers dominate latency in long-context decoding. Commodity GPUs with sixteen PCIe 4.0 lanes cap out around 32 GB/s, and pulling cache data across that link stalls the attention computation that's supposed to be running on the GPU at the same time. It's how to weave data-reduction techniques into a memory hierarchy that was never designed to hold this much state in the first place.
Limits of pure offloading and pure compression
KV reduction and KV offloading each address part of the problem but create new bottlenecks when applied without the other. Serving a large model at very long contexts requires KV cache memory far exceeding what a single 80 GB GPU can hold. It only delays the crossing point. Offloading the cache to slower, larger storage solves the capacity side, but it does nothing to slow the rate at which the cache grows in the first place, and stacking the two techniques together without care just trades one bottleneck for another: shrink too aggressively and accuracy suffers, offload too much and transfer volume drowns out any latency gained.
The tradeoffs are stark when tested in isolation. Systems that store the full cache in host memory and fetch only the pieces judged critical during decoding run into a ceiling of their own: sparsity can only be pushed so far before it starts costing accuracy, and there's no way to sparsify past that point without the model getting measurably worse. Worse, when context length and batch size grow at the same time, which is what production traffic does, the sheer volume of KV transfers spikes and becomes the single largest source of decoding latency.
Before layering anything more sophisticated on top, here is what problem is already solved. PagedAttention, the mechanism underneath vLLM, carves the KV cache into fixed-size blocks the way an operating system pages memory, eliminating most fragmentation waste and delivering a real throughput jump. That's the floor, not the ceiling. PagedAttention fixes how memory within a tier is packed; it says nothing about what happens when the tier itself runs out of room or how data should move between tiers once it does. Anyone building for million-token context windows has to start from PagedAttention's gains, but stopping there leaves the capacity and cross-tier transfer problems completely untouched.
What's missing across all of these partial fixes is coordination. CPU compute, GPU compute, data transfer, and storage all need to move in step, and existing approaches routinely fail at that, leaving some resource sitting idle while another one chokes. The actual engineering question, the one that tiered memory architecture exists to answer, is how to combine cache reduction and cache offloading so they reinforce each other instead of fighting over the same latency budget. GPU-only storage fails at long contexts, DRAM-only preserves capacity but suffers prohibitive latency, and uniform partitioning performs poorly on both metrics (per the research brief).
The four-tier memory hierarchy: what each layer offers and costs
Production systems serving long-context models now lean on four distinct hardware tiers: GPU HBM, host DRAM, local NVMe, and networked or disaggregated storage. Each tier trades bandwidth and latency for capacity and cost in a different way, and that tradeoff is what should decide what kind of KV data lives where.
GPU HBM sits at the top: highest bandwidth, lowest latency, and the least capacity of any tier by a wide margin. It's the only place the attention computation can read from without paying a transfer penalty, so it belongs to the KV states a request needs right now. Host DRAM comes next, bandwidth bound by whatever the PCIe link can carry, but offering far more capacity per dollar than HBM ever will. Host DRAM has substantially lower bandwidth than HBM but far greater capacity per dollar, making it appropriate for recently evicted KV states that may be needed within the current session.
Local NVMe extends the hierarchy further still. It trails DRAM badly on bandwidth but offers a large jump in raw capacity, and the field has largely converged on treating it as a required third tier for any serving system built around long context, not an optional extra. That said, NVMe only pays for itself when the GPU's KV cache is genuinely the bottleneck. Bolt it onto a short-context, single-user workload and it just adds I/O latency with nothing to show for it. Restoring KV cache from SSDs suffers from poor I/O performance: a fragmented GPU memory layout results in large numbers of tiny random I/Os, making the CPU a severe bottleneck even with GPU Direct Storage, which still requires CPU intervention to initiate each I/O.
Networked/disaggregated storage (RDMA-backed NVMe-oF, CXL memory pools, DPU-accelerated JBOF) offers TB-scale capacity, shared across compute nodes, enabling cross-node KV reuse. Large-scale serving systems that keep KV cache in local SSDs face fragmentation across nodes and limits on reuse when requests are rescheduled or migrated. Disaggregated storage sidesteps that by letting many compute nodes reach into one shared cache pool over a fast interconnect, and one such system reports up to a 35.7% throughput improvement over conventional CPU-offloading as a result.
It doesn't anymore. Agentic workloads and long-context serving have turned the cache into something closer to a persistent inference state, one that may need to survive across multiple turns and multiple sessions rather than disappearing when a single request completes. Modern inference systems rely on a multi-tier hierarchy spanning from capacity-constrained GPU memory to elastic remote storage.
Temporal relevance and tier placement for KV states
Handing out tiers by hardware property alone answers half the question. The other half is deciding which KV states go where, and the mistake most naive tiered systems make is treating every token in the cache as equally important regardless of when it entered. Existing approaches largely treat KV states as equally important across time, implicitly assuming uniform precision and accessibility.
That assumption doesn't hold up against how memory actually behaves, human or otherwise. This contrasts with how human memory systems work, where memories vary in clarity, recall frequency, and relevance with temporal proximity (the insight motivating TTKV). Recent tokens act like short-term memory, carrying most of the weight in what the model generates next, while older tokens fade into something closer to long-term memory, where only a small slice remains relevant to whatever the current query is actually asking. Temporal proximity, in other words, is the signal a placement system should be built around.
TTKV, a framework out of Harbin Institute of Technology and Guangzhou University, turns that principle into an explicit three-part design. Tier Layout splits the cache into a fast tier riding on HBM for short-term memory and a slow tier riding on DRAM for long-term memory, mapping the temporal split directly onto the hardware split. Tier Content then decides what precision each tier gets: recent, frequently accessed tokens keep full precision in the fast tier, while older and less frequently touched states get hit with differential quantization and sparsification once they land in the slow tier. Recent, frequently accessed tokens stay in HBM at full fidelity; everything older gets pushed to DRAM and compressed.
The last piece, Tier Interaction, is what keeps this from becoming a latency trap. Block-wise streaming attention overlaps computation with communication whenever the model needs to reach into the slow, DRAM-resident tier during decoding. That overlap is the entire reason the two-tier split works in practice, because without it every DRAM access would stall the GPU outright and cancel out whatever throughput was gained by tiering in the first place. On 128K-context tasks, this design cuts cross-tier traffic by a meaningful margin and delivers real gains in both latency and throughput over strong existing baselines. TTKV isn't the final answer to temporal placement, but it's the clearest existing proof that the principle holds up under actual measurement.
Predictive scheduling and pipeline overlap: eliminating stalls across all three tiers
Getting placement right solves the question of what belongs where. It doesn't solve when to move it. If a fetch from DRAM or NVMe blocks the GPU mid-computation, the throughput gains from correct tiering disappear regardless of how well the placement logic was designed. This is where the systems engineering gets genuinely hard.
Coordination failures occur in predictable places, caused by existing approaches that leave resources unsynchronized. GPU memory sits idle waiting on a transfer that hasn't landed yet, CPUs stall out during GPU execution windows, and I/O bandwidth goes unused in the gaps between transfers, all because existing approaches don't schedule CPU compute, GPU compute, data movement, and storage access as one system.
KVDrive, out of Hong Kong University of Science and Technology, takes on this stall problem directly and does it across all three tiers at once rather than optimizing any single link in isolation. Its attention-based cache management watches actual attention behavior to decide what to keep and what to move, maximizing reuse while cutting down on redundant transfers. Its scheduling layer decouples selection, fetching, and computation into separate stages that can overlap fine-grained, rather than running as one blocking sequence, which is what eliminates the stalls across GPU, DRAM, and SSD in the first place. And its storage layer coordinates movement across GPU memory, host DRAM, and SSD together, which is what lets long-context inference scale past what GPU and DRAM alone could ever support. The system achieves up to 1.74× higher throughput compared to existing systems while preserving accuracy. That gain isn't universal, though. The gain is workload-conditional: it helps when GPU KV cache is the actual bottleneck, while for short-context or single-user workloads it adds latency for no benefit.
TierKV, a framework built for mobile LLM inference, takes a different route to the same problem: prediction instead of reaction. Its core mechanism, Predictive Multi-Tier Cache Optimization, forecasts future cache demand from prefill hidden states before decoding even starts, then assigns tokens across exact, low-rank, and flash-offloaded tiers within whatever memory and accuracy budget the device has to work with. That's the meaningful architectural fork in the road: reacting to cache misses as they happen versus predicting demand ahead of time and paying a small prefill cost to avoid stalling mid-decode. Every decode step attends over the entire cached context, which makes each step its own potential I/O event, so scheduling has to be reasoned about at that per-step granularity rather than treated as a property of the request as a whole. SSD I/O fragmentation compounds this: fragmented GPU memory layout produces large numbers of tiny random I/Os, making CPU a severe bottleneck even with GDS (per the research brief).
Four production architectures for tiered KV management
The principles above, temporal placement, predictive scheduling, pipeline overlap, don't stay abstract for long once they hit production. TTKV shows one version of the tradeoff: a design operationalized into a three-part system, with block-wise streaming attention that overlaps computation and communication when accessing slow-tier (DRAM-resident) KV states during decoding. It's the leanest of the four architectures here, and that leanness is deliberate: two tiers, one clear signal (recency) governing both placement and precision.
KVDrive tackles GPU, host DRAM, and SSD stalls from a systems perspective across all three tiers, harmonizing data movement to unlock scalable long-context inference beyond GPU and DRAM limits. Its attention-based cache management adapts placement to attention behavior to maximize reuse and minimize redundant data movement. The tradeoff is complexity: three tiers and a scheduler coordinating all of them is a harder system to build and reason about than a two-tier temporal split, but it buys headroom that TTKV's design, by itself, doesn't reach for.
ITME is described as the architecture built for the reality that KV cache is no longer a transient, single-request artifact but something that may need to persist across multiple turns and sessions. ITME achieves up to 35.7% throughput improvement over conventional CPU-offloading, in a setting where keeping KV cache in local SSDs otherwise causes fragmentation across nodes and limits reuse when requests are rescheduled or migrated.
TierKV closes the set out at the opposite end of the deployment spectrum: mobile inference, where memory and accuracy budgets are set by the device itself, not by a data center rack. Its predictive approach, forecasting cache demand from prefill hidden states before decoding starts, is the clearest instance of the reaction-versus-prediction divide playing out in a real system, and it does it under constraints far tighter than anything the server-side architectures above have to contend with. Four systems, four different points on the same tradeoff curve: how much of the tier, precision, and scheduling problem gets solved ahead of time versus in the moment, and how far the hierarchy has to stretch, from a single GPU to an entire cluster, to keep serving the context a request actually needs.
Sources
- TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference
- ITME: Inference Tiered Memory Expansion with Disaggregated CXL-Hybrid Memories
- KVDrive: A Holistic Multi-Tier KV Cache Management System for Long-Context LLM Inference
- TierKV: Long-Context On-Device LLMs via Predictive Multi-Tier KV Caching — Large Language Models


