Est.

Multi-Tenant LLM Inference Serving and Cost Allocation

Token-based billing masks the true cost driver: GPU memory consumed by the KV cache.

Senior Editor, Infrastructure Economics · · 11 min read
Cover illustration for “Multi-Tenant LLM Inference Serving and Cost Allocation”
Scaling Economics · October 6, 2026 · 11 min read · 2,366 words

It's simple to quote, easy to compare across providers, and easy for a customer to forecast against a budget. That simplicity is also the problem: multi-tenant LLM platforms bill by the token, but the token is a poor stand-in for the resource that actually drives cost, which is GPU memory held by the KV cache over the lifetime of a request. Shared inference puts one GPU cluster to work serving many customers, with overhead spread across all of them, and that only works if effective cost per token stays roughly stable across tenants and requests. It doesn't, as on identical hardware, holding the model, precision, and GPU allocation fixed and varying only the offered request rate, effective cost per output token swings by a wide multiple within a single configuration, ranging from $0.21 to $15.25 per million output tokens across the load range tested. That spread has nothing to do with model size. It comes from concurrency and scheduling decisions made at the moment a request lands.

Consider two requests with the same token count on the invoice. One is short but drags along a lengthy system prompt that has never been seen by the server before; the other is longer but reuses a prefix the server cached minutes earlier. The first forces a full, uncached computation across every token in its context; the second reuses work already done. Measured in GPU-seconds, these two requests can differ enormously, yet a token-based invoice treats them as identical. That collapse means operators are pricing the wrong variable, and the error compounds as the platform scales: tenants who are cheap to serve get overcharged, tenants who are expensive to serve get subsidized by everyone else, and the operator loses the ability to explain its own margin swings from one month to the next.

Diagram: The Cost of a Token Varies by 72×. Visualizes: Show the dramatic spread in effective cost per output token across the load range tested on identical hardware, holding model, precision, and GPU allocation fixed.

How the KV cache binds GPU memory

Per-token billing fails for a physical reason, not just a statistical one. Every active request holds a chunk of GPU memory for as long as it runs, and that memory is consumed by the key-value cache, the stored record of past token representations that lets a transformer avoid recomputing attention over the entire sequence at each new generation step. The cache has to live in GPU high-bandwidth memory for the full duration of the request. So cost ties directly to memory occupancy over time, and that is a quantity per-token pricing never measures.

Supported context windows have grown from the tens of thousands of tokens common a few years ago to hundreds of thousands, and in some deployments millions. KV cache memory scales with context length, so every expansion in context window is, quietly, an expansion in the amount of GPU memory a single request can hold hostage. The arithmetic gets concrete fast. On an H100 serving a large model at long context, one user's KV cache runs roughly 40 GB. Eight concurrent users at that same context length need roughly 320 GB, which is four times the entire memory capacity of the GPU. Serving eight such users on one H100 requires moving cold KV blocks off the chip, so the storage architecture sitting behind the GPU is load-bearing infrastructure that the configuration depends on.

This reframes what "noisy neighbor" means in LLM serving. No tenant is stealing CPU cycles or saturating a network link from another. It's one tenant's KV cache blocks evicting another tenant's blocks out of a GPU memory pool that has a hard, fixed ceiling. Two tenants paying the identical per-token rate can have completely different experiences depending on who else is sharing that memory pool at that moment, and nothing in the billing system records which tenant caused which eviction.

KV cache allocation in serving engines

Underneath a multi-tenant platform, the serving engine's approach to KV cache allocation decides which workloads run cheaply and which run expensively, and that decision in turn decides which tenants the platform can serve at a profit. Two dominant approaches show this, and each bets differently on how workloads behave.

vLLM's PagedAttention splits the KV cache into small, fixed-size pages allocated on demand, rather than reserving a contiguous block of memory sized for the worst-case sequence length up front. That change cuts memory waste by up to 90% and lets far more requests fit into GPU memory at once, with reported throughput gains of up to 24 times over systems that pre-allocate. SGLang's RadixAttention takes a different bet: it organizes cached KV blocks in a radix tree keyed by token sequence, so that when two requests share a prefix, whether a system prompt, a retrieved document, or a set of few-shot examples, the attention computation for that shared prefix runs once and gets reused by every request that shares it.

Which bet pays off depends entirely on the tenant mix. On workloads with substantial prefix overlap, RadixAttention can deliver three to five times better effective prefill latency. But on workloads where requests share little or nothing, that advantage collapses to single-digit percentage differences. If a platform serves many tenants who all hit the same system prompt, it gains enormously from prefix deduplication, but if tenants send wildly different prompts with little structural overlap, it helps far less. This is a cost allocation decision as much as a performance one: choosing an engine means choosing which tenant profile the platform can serve economically, and neither engine makes the underlying capacity problem go away. When aggregate KV cache demand exceeds GPU HBM, something has to move, and the engine's eviction policy decides whose requests absorb the resulting latency.

KV cache overflow: the storage hierarchy and its costs

Diagram: The Storage Hierarchy Behind Every KV Cache Eviction. Visualizes: Illustrate the four-tier KV cache storage hierarchy and the bandwidth collapse between tiers: GPU HBM at ~3.35 TB/s (sub-microsecond latency), CPU DRAM at ~63 GB/s (tens of…

Once KV cache demand outgrows HBM, the system has two choices for every evicted block: send it to slower storage, or throw it away and recompute it later. Each option carries a cost set by the bandwidth and latency of wherever the block lands, and those costs can differ by orders of magnitude between tiers.

Production systems typically run a four-tier hierarchy: hot cache in GPU HBM, warm cache in CPU DRAM, cold cache on local NVMe, and shared or persistent cache on networked storage, coordinated by tiered caching systems such as LMCache. The bandwidth gaps between these tiers are stark. GPU HBM runs at roughly 3.35 TB/s with sub-microsecond latency. CPU DRAM, reached over PCIe 5.0 x16, runs at roughly 63 GB/s with latency in the tens of microseconds. Local NVMe SSD runs at roughly 7 GB/s, with latency from hundreds of microseconds up to low milliseconds. GPUDirect Storage-class paths let an NVMe reload bypass the CPU entirely, cutting the latency penalty of a cold reload but still leaving some penalty in place.

NVIDIA's Dynamo orchestration layer is built to operate across this entire hierarchy, from HBM through CPU DRAM to direct-attached NVMe and networked external storage, giving an operator one coordinated view of where KV state actually lives at any moment. Research on memory-expansion interconnects, described in work on ITME, proposes a further tier: terabyte-scale memory expansion reachable over standard RDMA, validated on production-grade memory-pooling hardware and NVMe hardware from another vendor, with feasibility shown through an FPGA-based prototype. So that tier matters if an operator needs capacity beyond what local NVMe can deliver cost-effectively.

The cost consequence runs straight through this hierarchy to the invoice, or rather, it should. If a tenant's requests consistently land in warm DRAM cache, they consume a meaningfully different, and cheaper, slice of infrastructure than a tenant whose requests are always cold and need a full NVMe reload. Per-token billing records none of this difference. Two tenants paying the identical rate per million tokens can be imposing entirely different loads on the storage hierarchy, and the invoice has no mechanism to tell them apart.

How disaggregated serving moves KV cache across the network

Large-scale serving has been moving toward disaggregated inference, where prefill and decode run on separate GPU pools. This is now the dominant architectural direction, and it solves real efficiency problems, but it comes at a structural cost: it turns the KV cache from a memory problem into a network problem, because the cache generated during prefill has to be transferred to wherever decode is running, for every single request.

The bandwidth available for that transfer depends entirely on the physical relationship between the sending and receiving GPU, and the range is wide. NVLink 4.0 moves data at 900 GB/s within a domain, but NVLink 5 pushes that to 1.8 TB/s. InfiniBand across nodes runs at roughly 50 GB/s. TCP across datacenters runs at roughly 12.5 GB/s. So you get a wide gap between the fastest and slowest path under NVLink 4.0, and an even larger gap once NVLink 5 is in the mix.

Most disaggregated serving systems in production do not account for this spread. DistServe uses uniform RDMA regardless of where the two endpoints sit in the topology. Splitwise sidesteps the problem by co-locating prefill and decode on the same machine. Mooncake uses a Transfer Engine that supports both RDMA and TCP with some topology awareness, but systems built this way still end up paying NVLink-class latency for some transfers and TCP-class latency for others, without the cost difference being tracked or billed anywhere. Two tenants submitting identical requests can end up with very different latency and very different infrastructure cost, purely because of where their prefill and decode slots happened to land in the cluster. Per-token billing has no way to represent that outcome, because it never measured the network hop that caused it.

KV cache compression as a lever that changes what multi-tenant economics are possible

Compression changes what you can do economically in multi-tenant serving; it doesn't just make an existing configuration run faster. Shrinking the KV cache stretches how much GPU memory capacity you actually have, cuts how much data has to move across the storage tiers described above, and shifts the concurrency level at which a given hardware footprint stays profitable. A 2026 survey of KV cache optimization techniques groups the field into five categories, and in production, what matters most is pairing memory layout improvements with lossy or lossless compression of the cached values themselves.

On the lossy side, Google's TurboQuant, announced in March 2026, compresses KV cache values down to 3 bits per value with no measured loss in accuracy, cutting KV cache memory by 6 times. NVIDIA's KVTC, presented at ICLR 2026, claims it can compress up to 20 times using transform coding while reasoning and long-context accuracy hold steady.

On the lossless side, research on Huffman-based compression of BF16 model weights has found that compressibility in BF16 shifts from the mantissa to the exponent bits, meaning models that compress poorly in FP16 can compress well in BF16 because of the larger exponent field. The DFloat11 approach exploits the uneven statistical distribution of exponent bits found in trained LLM weights, so it can losslessly compress down to 70% of the original BF16 model size with no accuracy loss. SplitZip targets a narrower but directly relevant case: ultra-fast lossless KV compression built specifically for disaggregated serving, reducing the volume of KV state that has to cross the network between prefill and decode pools described in the previous section. BF16 itself is now the default precision for LLM inference, natively supported across NVIDIA Tensor Cores, Google TPUs, and Intel AMX, making BF16-specific compression research a production concern.

The cost allocation consequence follows directly. If a tenant's requests produce short, highly compressible KV blocks, that tenant occupies less effective memory capacity than one whose requests generate large, incompressible blocks, even when both send the same number of tokens. That difference is invisible on a per-token invoice. It is fully visible, and fully measurable, in memory-hours consumed on the GPU.

The concentration problem: why a small fraction of tenants drive most KV cache cost

KV cache cost does not spread evenly across a tenant base. Long-context power users, RAG pipelines pulling in large documents, and reasoning-heavy workloads occupy disproportionate amounts of GPU memory for disproportionate stretches of time. This concentration is a structural feature of how these workloads use context, not an artifact of any particular customer's behavior.

Operators without per-tenant KV cache attribution end up optimizing the wrong thing. Aggregate throughput looks fine on a dashboard even while a handful of tenants with enormous context windows are filling the shared memory pool and forcing evictions that degrade service for everyone else sharing that GPU. A long prompt costs far more to serve uncached than cached, and time to first token runs several times higher on a cold cache hit than a warm one. If a tenant's requests reliably hit warm cache, that tenant consumes a fraction of the GPU time of a tenant whose requests are always cold, even if both are billed the identical rate per token. Left unmeasured, that gap doesn't average out. It sits on the operator's margin as unexplained variance, tenant by tenant, until someone goes looking for where the GPU memory actually went.

Token pools: representing KV cache capacity as an explicit entitlement

The token pools abstraction, introduced by William Cunningham of DataRobot at the ACM Conference on AI and Agentic Systems in May 2026, responds directly to this attribution failure. Rather than treating multi-tenant capacity as a rate limit, token pools represent it as an explicit entitlement defined in the units that actually determine cost: token throughput, KV cache, and concurrency.

The distinction from conventional rate limiting is specific. A rate limit governs whether a request is admitted, without any regard for what that request will cost to execute once it's running. Token pools authorize both admission and autoscaling from the same underlying capacity model, which keeps what a platform promises a tenant consistent with what it actually provisions for that tenant. The framework defines four service classes, each carrying different guarantees. A dedicated class reserves capacity for a tenant but allows that tenant to burst onto idle capacity belonging to others when it's available. A guaranteed class reserves capacity as well, but without any right to burst beyond it. These classes give an operator a way to sell what the infrastructure actually delivers, rather than selling a token count that was never the resource in short supply to begin with.

Sources

  1. Token Management in Multi-Tenant AI Inference Platforms
  2. Beyond Per-Token Pricing: A Concurrency-Aware Methodology for LLM Infrastructure Cost Estimation

More in Scaling Economics