Power and Cooling Costs in GPU Inference Data Centers
Power and cooling infrastructure now costs as much as the GPUs themselves.

Power, not GPU supply, is now the thing standing between an inference workload and production. Through 2023 and into 2024, the industry's bottleneck was silicon: could you get allocation on H100s, could you get them fast enough. By 2026, the constraint has moved downstream, past the chip and into the power infrastructure. Getting the GPUs is often the easy part now. Getting the megawatts to run them, cool them, and keep them running is where projects stall.
Power as the binding constraint on GPU inference at scale
The shift raises grid demand, as reflected in the numbers. The IEA's 2025 "Energy and AI" report projects that global data center electricity consumption will double by 2030, and AI workloads are widely cited as a primary driver of that incremental demand, not general-purpose computing. That's a demand curve utilities and grid operators were not planning for a decade ago, and it shows in how long it now takes to bring new capacity online.
In Northern Virginia, in Silicon Valley, and across Northern Europe, power approval timelines for new facilities have stretched substantially, regardless of whether the hardware sitting in a warehouse is ready to plug in the same week. A developer can have GPUs on a loading dock and still wait two years for the utility interconnect that lets those GPUs draw current.
The power draw per rack has moved so fast that the planning assumptions from three years ago no longer apply. Rack power has climbed from roughly 25 kW to more than 130 kW in about two years, and the trajectory points past 600 kW for the densest configurations coming next. Facilities designed around 2023 assumptions are not just a little short. They are off by an order of magnitude, and no amount of clever scheduling fixes a building that was never wired for this.
GPU hardware choices and the floor they set for power and cooling requirements
The GPU's thermal design power (TDP) number is not a ceiling that gets touched occasionally under a spike. Under sustained inference load, that number is close to the continuous draw the hardware pulls, hour after hour. Procurement teams who treat TDP as a worst-case figure end up under-provisioning power and cooling from day one.
The GPU generation on the purchase order sets a floor for facility requirements, and that floor varies enormously by SKU. NVIDIA L40S inference servers run at a comparatively modest per-GPU TDP and are 30 to 40 kW per rack, air cooled, workable in most tier 3+ colocation space, and a reasonable fit for 7B to 13B parameter models. Step up to H100 and H200 servers and rack draw runs 40 to 120 kW, the range where most mainstream 13B to 70B production inference lives today. AMD's Instinct MI300X and MI325X are in a similar 60 to 100 kW band, air or direct liquid cooled, and their large memory footprint makes them well suited for large-parameter models.
NVIDIA's B200 Blackwell servers push further: 1,000W per GPU, 60 to 150 kW per rack, with an 8-GPU server alone drawing approximately 14.3 kW. And at the top end, the GB200 NVL72 and its successors (GB300, Rubin) run 130 to 600-plus kW per rack and require direct-to-chip liquid cooling as a condition of deployment, not an upgrade option. Only facilities carrying DGX-Ready certification can host them.
None of that accounts for the rest of the server. On an 8x H100 node, the GPUs themselves make up only just over half of total node power draw. CPUs, NVLink switches, RAM, and power supplies fill in the rest, and the real number is around 10 kW per node, not the 5.6 kW a back-of-envelope calculation (8 times 700W) would suggest. Teams that size power budgets off GPU TDP alone are routinely undershooting by close to half.
The practical consequence is a filtering effect on the colocation market. The highest-density configurations, a top-tier class of racks, can only be hosted in a small slice of global colocation capacity, a small slice of global colocation capacity, because most buildings simply were not built to deliver that much power to a single rack, let alone remove that much heat from it.
Facility build and operational electricity costs for inference operators
Building a facility to house this hardware costs meaningfully more than a conventional hyperscale build. Industry figures put a data center built for this hardware at a cost that runs several times what a conventional facility costs, with the premium especially pronounced at the high end. That premium buys the electrical infrastructure and cooling plant that GPU racks require, and it explains why so few operators can simply repurpose existing space.
Existing space, for the most part, cannot absorb this load anyway. The industry mean rack density across conventional facilities is around 7.6 kW, which is a rounding error next to a 130 kW GPU rack. A facility built for enterprise virtualization workloads requires a substantial retrofit to host inference at scale, if it can be retrofitted.
Then there's the electricity bill itself, and geography swings it wildly. For an identical 1,000-GPU cluster (a continuous draw north of 1 MW), monthly electricity cost ranges from roughly $50,800 to $317,500 depending on where the facility sits. That's better than a sixfold difference, driven entirely by local $/kWh rates, not by anything about the hardware or the workload. Two operators running identical clusters can have wildly different unit economics purely because of where they signed a lease.
At the smaller end of the market, a mid-market colocation deployment, something like 4 racks at 30 kW each, runs somewhere in the neighborhood of the low tens of thousands of dollars monthly once power, cooling, and amortized hardware over a 4-year refresh cycle are all counted. That's a useful anchor for teams sizing a first deployment: the electricity line item is not an afterthought next to the hardware lease. It's a comparable cost.
The real mechanics of PUE: how cooling architecture multiplies or reduces every watt of compute spend
Power Usage Effectiveness (PUE) is the ratio that tells an operator how much of the electricity coming into a building actually reaches the servers, versus how much gets spent moving heat back out. Industry average PUE in 2026 is around 1.3 to 1.5, and 30 to 50% of total facility power goes to cooling, lighting, and other overhead rather than compute. Legacy air-cooled facilities fare worse, stuck around 1.55 to 1.67, which means something close to 40% of total electricity spend buys nothing but heat removal.
The best hyperscale facilities do far better. PUEs below 1.1 are achievable, and liquid-cooled facilities routinely run in the 1.04 to 1.1 range. That gap between 1.5 and 1.1 amounts to paying for compute once versus paying for it one and a half times over. It's the difference between paying for compute once and paying for it one and a half times over.
A concrete example makes the arithmetic clear. A 10 MW AI facility running at PUE 1.5 pulls 15 MW total from the grid, with 5 MW of that lost purely to cooling overhead. Move that same facility to direct-to-chip cooling at PUE 1.10, and total draw falls to 11 MW, a meaningful reduction in annual energy costs that can change whether a deployment is profitable. It's a line item large enough to change whether a deployment is profitable.
PUE measures how much power reaches the IT load, but it says nothing about what that IT load actually produces. A facility can have an excellent PUE and still burn power on GPUs sitting execution-idle (more on that below), producing no tokens. Water Usage Effectiveness (WUE) and similar metrics are starting to gain traction precisely because operators need a way to ask whether the power delivered is turning into useful compute, not just whether it's being delivered efficiently.
Cooling technology options and what each one delivers at GPU inference densities
Air cooling with rear-door heat exchangers remains viable for lower-density racks, the L40S class at 30 to 40 kW, and it works in most tier 3+ facilities without special retrofits. But it hits a physical wall at higher rack densities. Past that point, air simply cannot move enough heat fast enough, no matter how the airflow is engineered.
Direct-to-chip, or cold plate, liquid cooling circulates coolant directly over the GPU die and enables densities of 80 to 120 kW per rack, with PUE typically in the 1.04 to 1.1 range. A growing number of facilities are shifting toward supplying hot water, in the range of 18°C to 25°C, rather than chilled water, because it cuts the energy the chiller plant itself consumes. Warmer inlet water reduces cooling energy spend, but it narrows the thermal margin on the silicon, and under bursty load it can produce temperature spikes that a colder loop would have absorbed.
Immersion cooling goes a step further by submerging hardware directly in a dielectric fluid. Single-phase immersion, where the fluid stays liquid throughout, delivers PUE in the 1.04 to 1.08 range with comparatively low deployment complexity. Two-phase immersion, where a low-boiling-point fluid vaporizes on contact with hot components and carries heat away as it changes state, pushes PUE down further still, into the 1.02 to 1.05 range, with the best implementations reaching 1.01. The trade-off is complexity and cost: two-phase systems require more specialized fluid handling and carry a higher price tag than single-phase, which is part of why adoption has lagged behind the efficiency numbers alone would suggest.
Execution-idle: the hidden power state that makes GPU utilization the most important cost lever
Here's a finding that runs against intuition: a GPU can sit at high power draw even while doing, visibly, almost nothing. On a CPU, power tracks activity closely, idle cores drop to a low baseline almost immediately. GPUs don't behave that way. A GPU can have a program loaded and memory allocated, with compute, memory, and communication activity all near zero, and still pull power close to its active load.
This state has a name: execution-idle. It describes the interval where the GPU remains allocated and a kernel or program stays resident, but there's no meaningful work happening on it, distinct from a truly idle GPU that has been released and drops back to a low baseline. The distinction matters because execution-idle GPUs look busy to a scheduler (they're allocated, they're "in use") while contributing nothing to throughput.
A 31-day study of a large academic AI cluster, spanning six GPU platforms including B200, measured just how large this gap is: execution-idle accounted for 19.7% of in-execution time and 10.7% of total energy consumed across a range of workloads. That's real money spent on GPUs that are neither idle enough to power down nor active enough to produce anything.
Serving workloads make the problem worse, not better. Bursty request arrival patterns, the normal shape of production inference traffic, create loaded-but-inactive gaps between requests, where the GPU sits allocated to a serving process but waits for the next token request to arrive. In long-lived academic serving workloads, execution-idle accounted for 48% of total energy. Replays of industry-derived traces from OpenAI, Qwen, and Azure in the same research (Lei et al.) showed execution-idle ranging from 7% to 65% of energy depending on the traffic pattern. That range alone functions as the single largest lever most operators haven't pulled yet, not a soft efficiency metric, and it should reframe how inference teams think about utilization.
KV cache architecture's connection between memory pressure and power cost in production inference
Every transformer-based inference request builds a KV cache, the stored keys and values that represent relationships between tokens as a sequence grows. The longer the context, the bigger the cache, and it grows in a way that's awkward for fixed GPU memory: past a certain context length, KV cache routinely exceeds what fits in HBM. When that happens, the system either evicts and recomputes (burning GPU cycles and power to redo work it already did once) or offloads to another memory tier (burning interconnect and storage energy instead).
Memory fragmentation compounds the problem quietly. Traditional inference systems waste a large majority of allocated KV cache memory to fragmentation and over-allocation, and the GPU holds far more memory reserved than it's actually using productively at any given moment. That wasted allocation does double damage: it tightens memory pressure, forcing earlier eviction or offloading, and it's part of what produces the execution-idle conditions described above, GPUs sitting allocated and holding memory while doing comparatively little useful work.
Systems handling this well use a four-tier hierarchy: hot KV cache in GPU HBM, warm cache in CPU DRAM, cold storage on local NVMe, and shared or persistent storage across the network for cache that needs to survive beyond a single request. Frameworks like NVIDIA Dynamo and LMCache exist specifically to manage movement across these tiers. Every tier transition costs something, though: PCIe or RDMA transfer energy, and latency that can stretch out the execution-idle window while the GPU waits for data to arrive from a slower tier.
A 70B parameter model running an 8K context window requires roughly 20 GB of KV cache, and at batch scale, across many concurrent requests, KV cache consumption often exceeds the memory footprint of the model weights themselves. Storage architecture is the primary factor determining how many requests a GPU can serve concurrently before offloading becomes unavoidable. It's the primary factor determining how many requests a GPU can serve concurrently before offloading becomes unavoidable.
The electricity economics of older GPU hardware: when second-hand infrastructure changes the cost calculus
Large AI data centers retire functional GPUs on roughly three-year cycles to make room for the next generation, and those retired GPUs don't stop working, they move to secondary markets at a fraction of original cost. The price gap is stark: a 128-GPU cluster assembled entirely from second-hand components cost approximately $22,000, against roughly $600,000 for a current-generation 8-GPU B200 system. It's a different economic category.
The question that price gap raises is whether older, cheaper, less power-efficient hardware can still make sense once electricity costs are factored in, and one documented case suggests the answer is yes under the right conditions. One project built and ran a 128-GPU cluster from second-hand units of an older-generation chip over the course of a year, and using pipeline-parallel optimizations, it reached competitive LLaMA-70B throughput despite running on hardware several generations behind the frontier. That older chip draws far less per-chip performance than the newer chips that succeeded it, and per-token energy cost on older silicon is correspondingly higher. But when the capital cost gap runs into the tens of multiples, a slower, cheaper, less efficient cluster can still win on total cost per token for workloads that tolerate the throughput trade-off.
That's the calculation every inference operator eventually has to run for themselves: not which GPU is fastest, but which combination of capital cost, electricity price, cooling overhead, and utilization gets the lowest cost per token for the workload actually being served. Power and cooling are variables inside that decision, and the operators treating them that way are the ones who'll have inference economics that hold up as density keeps climbing. They're variables inside it, and the operators treating them that way are the ones who'll have inference economics that hold up as density keeps climbing.
Sources
- AI Inference Power Consumption and GPU Electricity Costs: 2026 Guide | Spheron Blog
- AI Inference Colocation: Power, Cooling, and Network Requirements for GPU-Ready Data Centers
- The Energy Cost of Execution-Idle in GPU Clusters
- DumpsterCluster: From Dumpster Diving to Serving LLaMA-70B on 60 GPUs
- The Model Parking Tax: Quantifying the Hidden Energy Cost of Always-On GPU Model Deployment
- AI Data Center Cost Per MW: Budgeting a 2026 GPU Buildout | Gain America
- mlq.ai
- AI Data Centers Energy Consumption in 2024–2026: Trends, Projections, Environmental Impact Investment Opportunities | TTMS


