What $6.69 Per GPU-hour Really Costs

AI Datacenter advisory · September 4, 2026 · 7 min read

A rate card is the easiest number in this industry to read and the hardest to interpret. $6.69 per GPU-hour for an on-demand B200 looks like a clean comparison against owning the same GPU. It is not, because the two sides of the comparison hide different things.

What the Hourly Price Includes

The neocloud price bundles the GPU, the server, the rack, the power, the cooling, the fabric, a share of storage, the building, staff, financing and margin. On-demand carries the premium for flexibility; reserved and committed contracts drop the price by 40–60% in exchange for a one- to three-year commitment, which is itself a kind of ownership without the asset at the end.

What it does not include is egress, high-performance storage beyond the baseline, support tiers, and the cost of your team's time when allocation is tight and jobs wait.

What Owning Includes

Owning an NVL72 rack means capital for the rack and its share of facility, power and cooling plant; electricity at your tariff; maintenance and support contracts; staff or a managed NOC; and the refresh question at year three to five. Against that, there is no hourly meter, and the asset has a residual value that neocloud spend never does.

The owned case lives or dies on three variables. Utilisation: an owned GPU costs the same whether it runs or idles, so a cluster at 40% utilisation is paying more than double per useful hour than one at 85%. Power price: at 132 kW per rack, each cent per kWh is roughly $11,500 a year per rack. Refresh cycle: a three-year refresh roughly doubles the annualised capital cost versus five years.

Where the Break-Even Tends to Land

On the models we have run this year, a steady training or inference load above about 60–65% utilisation on committed pricing, or above 40–45% against on-demand pricing, breaks even on ownership inside the second year and is materially cheaper over five. A bursty workload that runs hard for six weeks a quarter does not, and should stay rented — or run a small owned baseline with rented burst.

The point is not that owning wins. It is that the answer depends on your utilisation curve and your power tariff far more than on the GPU price, and those are the two numbers most teams have not measured when they come to us.

A Fair Comparison

Model the same five years on both sides, with the same utilisation assumptions, the same storage and egress, and a realistic operating cost for the owned case. Then stress it: what happens if utilisation is 20 points lower than planned, if power rises two cents, if the next GPU generation makes a three-year refresh attractive. If ownership only wins in the base case, rent. If it wins across the range, build. Most honest analyses end up with a hybrid, which is why we designed our advisory engagement to be independent of whether a build follows.

Related service

Build-vs-Rent Advisory

Independent TCO against live neocloud rate cards.

More from the Blog