Skip to content

AI Rack Cost: How to Price Accelerated Compute End to End

Capital, energy times PUE, and overhead, divided by capacity times utilisation. The framework, the multipliers that matter, and the five places the arithmetic reliably breaks.

Noorain Fathima · 11 min read
A liquid-cooled server rack beside a cost breakdown showing capital, energy and utilisation as separate contributions
A liquid-cooled server rack beside a cost breakdown showing capital, energy and utilisation as separate contributions
Contents
  1. The quick answer
  2. Key takeaways
  3. The cost equation
  4. Capital: what a rack now contains
  5. Power, and the multiplier you cannot avoid
  6. Cooling is a facility decision, not a line item
  7. Utilisation: the term that decides the answer
  8. Working it through
  9. What the demand trend does to the inputs
  10. Where the arithmetic usually breaks
  11. Frequently asked questions
  12. What is a good PUE for an AI data centre?
  13. Is it cheaper to buy GPUs or use an API?
  14. Why is the rack the unit of purchase rather than the GPU?
  15. How should I estimate tokens per hour?
  16. How long should I amortise an AI rack over?
  17. Final takeaway
  18. Sources and further reading

Vendors quote peak FLOPS. Operators pay for power, floor space and the hours the machine sits idle. AI rack cost is the arithmetic that connects the two, and it is the calculation most often skipped before a purchase order is signed — usually because the inputs live in four different departments.

The structure is not complicated. Capital cost amortised over a life, plus energy at the facility's real multiplier, divided by the fraction of capacity you actually use. What makes it hard is that the third term is the one nobody wants to estimate honestly, and it moves the answer more than the first two combined.

The quick answer

Effective cost per unit of work = (amortised capital + energy × PUE + facility and staffing overhead) ÷ (nameplate capacity × utilisation). Capital dominates the numerator for accelerated compute, energy is the fastest-growing term, and utilisation is the term with the widest range — a rack at 30% utilisation costs more than three times per token what the same rack costs at 100%. Get a real utilisation estimate before comparing against an API price, because that is the comparison the number is usually being used to make.

Key takeaways

  • Cost per token is a function of four inputs, and only one of them appears on the vendor's datasheet.
  • PUE is a multiplier on every watt: at Google's published fleet average of 1.09 you pay 9% overhead; at the industry average of 1.54 you pay 54%.
  • Utilisation is the dominant uncertainty. Halving it doubles your cost per token with no other change.
  • Accelerated server electricity consumption is projected by the IEA to grow 30% a year, against 9% for conventional servers.
  • Do not use vendor list prices in this calculation. Use your quote, because large-scale accelerator pricing is negotiated and not public.
  • The rack is the unit that matters now, not the GPU — memory and interconnect are pooled at rack scale.

The cost equation

Write it out before filling anything in:

TermWhat it isWhere to get it
Capital, amortisedPurchase price ÷ useful life in hoursYour negotiated quote; life is a finance decision
EnergyRack draw in kW × hours × price per kWhFacility spec and your power contract
PUE multiplierTotal facility power ÷ IT powerYour colocation provider, measured not marketed
OverheadSpace, networking, staff, support contractsUsually the most-underestimated line
Nameplate capacityTokens per hour at your batch and precisionMeasured on your workload, not a benchmark
UtilisationFraction of that capacity actually usedYour traffic pattern — estimate honestly

Two of these are facts you can look up. Four require work. That asymmetry is why the calculation so often gets truncated into "price divided by FLOPS", which answers a question nobody is actually asking.

Capital: what a rack now contains

The meaningful unit has shifted from the accelerator to the rack, because memory and interconnect are now pooled across the whole enclosure rather than isolated per server.

NVIDIA's GB200 NVL72 makes the shift concrete. One rack contains "36 Grace CPU | 72 Blackwell GPUs", "13.4 TB HBM3E | 576 TB/s" of GPU memory, "17 TB LPDDR5X" of CPU memory, and "130 TB/s" of NVLink bandwidth linking the GPUs into a single domain. That last figure is the reason the rack is the unit: 72 GPUs connected at 130 TB/s behave, for a large model, like one very large accelerator. The same 72 GPUs in separate chassis on a slower fabric do not.

What NVIDIA does not publish on that page is a rack power figure, and you need one. Get it from the vendor quote or the facility engineering specification, because everything downstream in the energy term depends on it, and because rack densities at this class have moved far enough that older facility assumptions frequently do not hold.

On price: this article does not print one. Accelerator pricing at scale is negotiated, varies by volume, region and bundled support, and any figure published here would be wrong for your situation and stale within a quarter. Use your own quote. The framework is what generalises; the number is not.

Power, and the multiplier you cannot avoid

Every watt reaching a GPU brings overhead with it — cooling, power conversion, distribution losses. PUE is the ratio, defined as total facility energy divided by IT equipment energy, and it multiplies your entire energy bill.

The spread between good and typical is much larger than most cost models assume. Google reports that "in 2025, the average annual power usage effectiveness for our global fleet of data centers was 1.09", measured comprehensively — a "trailing twelve-month (TTM) PUE ... across all our large-scale data centers (once they reach stable operations), in all seasons, including all sources of overhead". Against that, Google cites the Uptime Institute's 2025 Global Data Center Survey, in which "the global average PUE of respondents' data centers was 1.54".

At 1.09 you pay a 9% energy premium for the building. At 1.54 you pay 54%. On a large accelerated deployment that difference is not a rounding error — it is a line item that can exceed the salary cost of the team operating it.

Two cautions when someone quotes you a PUE. Ask what it includes and over what period: a design PUE at full load in winter is a different animal from a measured trailing-twelve-month figure that includes partial-load operation and summer cooling. And ask whether it is measured at all, or modelled. Google's phrasing — "including all sources of overhead" — is doing deliberate work, because the common way to report a flattering PUE is to exclude some.

UniverseBlend's survey of the hidden limits on AI compute covers the facility-side constraints that sit behind these numbers, several of which bind before the electricity price does.

Cooling is a facility decision, not a line item

At current rack densities, cooling stops being something the building does in the background and becomes a constraint on whether the deployment is possible at all.

NVIDIA describes the GB200 NVL72 as a liquid-cooled, rack-scale design, and that is not a preference — air cooling does not remove heat fast enough at these densities to be practical. The consequence for costing is that a facility built for conventional server loads may not be able to host the rack without modification, and the modification is a capital project with its own timeline.

Three questions belong in the cost model before the hardware is ordered. Does the target facility support direct liquid cooling, and if not, what does retrofitting cost and how long does it take? What is the power density limit per rack, and does it accommodate the draw the vendor specifies? And what does the cooling approach do to PUE — liquid cooling generally improves it relative to air at high density, which is a genuine offset against the retrofit cost, but the size of that offset is site-specific.

Teams that skip this discover it at the worst possible moment: hardware delivered, facility unable to accept it, and a lead time on the remedy measured in quarters. The cheapest time to find out is during procurement, when it is still a question rather than a problem.

Utilisation: the term that decides the answer

Capital and energy are reasonably knowable. Utilisation is where cost models quietly fail, and it fails in a predictable direction: optimistically.

The arithmetic is unforgiving because utilisation sits in the denominator. A rack costing a fixed amount per hour, running at 100% of its useful capacity, produces some cost per token. The same rack at 50% produces double. At 30%, triple and a third. Nothing about the hardware changed.

UtilisationEffective cost multiplierWhat this typically looks like
90%1.1×Batch workloads with a deep queue
60%1.7×Well-managed serving with mixed workloads
30%3.3×Interactive serving sized for peak traffic
15%6.7×Dedicated capacity for a single bursty application

Interactive serving is structurally prone to the lower rows. You size for peak, traffic is diurnal, and the capacity you provisioned for the busiest hour is idle for most of the day. Reserved cloud capacity has the same property; the only difference is who owns the idle hardware.

This is also where the memory-bandwidth analysis bites. As covered in why memory bandwidth beats peak FLOPS, decode only approaches the hardware's arithmetic capability at large batch sizes — so a deployment with insufficient concurrency is under-utilising the silicon even when the GPU appears busy. "Utilisation" measured as GPU occupancy and utilisation measured as useful work per dollar can differ by an order of magnitude.

Working it through

The method, with placeholder inputs you replace with your own — none of the numbers below are researched market figures, and they are here only to show the shape of the calculation.

Suppose a rack costs C and you amortise it over three years of continuous operation, which is 26,280 hours. Suppose it draws P kilowatts, your facility runs at PUE 1.3, and you pay E per kilowatt-hour. Then hourly cost is:

(C ÷ 26,280) + (P × 1.3 × E) + overhead

Measure your own tokens per hour at your batch size, precision and context length — not from a benchmark, because the prefill-to-decode ratio of your traffic changes the answer substantially. Multiply by your honest utilisation. Divide.

If the model you intend to serve is an open-weight one, the memory arithmetic in what open-weight models cost to hold in HBM decides how many racks the deployment needs before any of this applies — a model that does not fit is not a cost question yet.

Now compare against an API price for the same model quality. Three things usually emerge. Self-hosting looks better at high sustained utilisation and worse at low. The crossover point is higher than enthusiasm predicts. And the comparison is only valid if you have included overhead, which teams routinely omit — the networking, the support contract, the engineers who keep the serving stack running. UniverseBlend's breakdown of the hidden fees in a self-hosted bill is a useful checklist for the lines that go missing.

What the demand trend does to the inputs

Two of the terms above are moving, and in the same direction.

The International Energy Agency estimates that "electricity consumption from data centres is estimated to amount to around 415 terawatt hours (TWh), or about 1.5% of global electricity consumption in 2024", projected to reach "around 945 TWh by 2030" — "just under 3% of total global electricity demand". Data centre consumption "has grown by around 12% per year since 2017, more than four times faster than the rate of total electricity consumption".

The composition matters more than the total for anyone costing accelerated compute. The IEA projects electricity consumption in accelerated servers growing "by 30% annually in the Base Case, while conventional server electricity consumption growth is slower at 9% per year", with accelerated servers accounting for "almost half of the net increase in global data centre electricity consumption".

The way those pressures reach a published price per token is the subject of why token pricing is a facilities question. What follows is not a price prediction — the IEA does not make one and neither will this article. What follows is a structural observation: demand growing at 30% a year against generation and transmission capacity that expands on multi-year timelines puts upward pressure on the energy term and on the availability of sites with power at all. Model your energy input as a range rather than a point, and treat a long-term power contract as a genuine asset rather than a procurement formality.

Where the arithmetic usually breaks

Five recurring failures, in rough order of how much damage they do.

  • Optimistic utilisation. The single largest error. If you cannot defend the number, model the pessimistic case and see whether the decision still holds.
  • Omitted overhead. Networking, storage, support contracts, and the engineering time to operate a serving stack. This is frequently 20–30% of the total and is left out entirely.
  • Design PUE instead of measured. A modelled figure at full load flatters the energy term by a margin that compounds over the life of the deployment.
  • Benchmark throughput instead of workload throughput. Your prompt and output length distribution determines the prefill-to-decode ratio, and therefore the tokens per hour the hardware actually delivers.
  • Amortisation longer than the useful life. Accelerator generations arrive faster than three-year finance cycles assume, and a rack that is no longer competitive is a cost, not an asset.

Frequently asked questions

What is a good PUE for an AI data centre?

Google publishes a comprehensive trailing-twelve-month fleet average of 1.09 for 2025, which represents the well-optimised end of hyperscale operation. Google cites the Uptime Institute's 2025 survey figure of 1.54 as the global average of respondents. Anything you are quoted should be measured over a full year including all overhead, not modelled at design load.

Is it cheaper to buy GPUs or use an API?

It depends almost entirely on utilisation. Owned hardware has a fixed cost that is incurred whether or not you use it; an API charges per token. Above a sustained utilisation threshold, ownership wins; below it, the idle capacity dominates. Calculate your own crossover rather than adopting someone else's, because the threshold is sensitive to inputs specific to you.

Why is the rack the unit of purchase rather than the GPU?

Because memory and interconnect are pooled at rack scale. A GB200 NVL72 links 72 GPUs with 130 TB/s of NVLink bandwidth and presents 13.4 TB of HBM3E, which lets a large model be sharded across the rack and behave as though it were on one very large accelerator. The same GPUs on a slower fabric behave like separate machines.

How should I estimate tokens per hour?

Measure it on your own traffic. Published benchmark numbers use fixed prompt and output lengths that may not resemble your workload, and the prefill-to-decode ratio materially changes throughput on identical hardware. Run your actual request distribution at your intended batch size and precision.

How long should I amortise an AI rack over?

Shorter than a general-purpose server, and the reason is competitive rather than physical. The hardware keeps working; the question is whether it remains the cheapest way to serve a token once the next generation ships. Model the sensitivity — if the decision only works at a five-year life, it is a fragile decision.

Final takeaway

The cost of accelerated compute is not a property of the hardware. It is a property of the hardware, the building it sits in, the power contract behind it, and — most of all — how much of the time it is doing useful work.

Three of those four are outside the vendor's control and outside the datasheet. Do the division yourself: capital plus energy times PUE plus overhead, over capacity times utilisation. If the answer only works at utilisation you cannot defend, you have learned the most valuable thing the calculation had to tell you, and you have learned it before signing.

Sources and further reading

0 likes, 0 saves

Found this useful? It helps to know.

Written by Noorain Fathima

AI engineer specialising in agentic systems and founder of MJ Smart Solutions in Bengaluru, building intelligent document processing, voice assistants and multi-agent platforms. Writes the Nexus on compute economics, model governance and agent security. Writing since March 2026. A published researcher and a product and UI/UX designer as well as an engineer, and studied at REVA University. That mix is the standard the Nexus holds itself to: sources opened and read rather than summarised second-hand, figures checked against the footnotes they come from, and every outbound link verified before a piece publishes.

Noorain Fathima on LinkedIn

Comments

No comments yet. Corrections and disagreements are especially welcome.

Leave a comment

Not published. Used only so we can reply.

Comments are reviewed before they appear.

Read Next

See all
A token price list overlaid on an electricity transmission tower, linking inference pricing to power infrastructure

LLM Token Pricing Is a Facilities Question

Emerging Technology

Output costs five times input because decode is memory-bound. Batch costs half because utilisation is the provider's central problem. Beneath both sits a floor made of power contracts and grid queues.

10 Sept 2026 · 12 min read

Subscribe to our newsletter

Occasional dispatches on AI, robotics and the engineering behind them. No spam, unsubscribe in one click.