Below the Neocloud Floor: the 2 MW Private Cluster
AI Datacenter engineering · August 21, 2026 · 6 min read
The large GPU clouds are optimised for very large tenants. A two-year, 100 MW commitment gets attention and allocation; a request for 16 racks in your own building does not fit their business at all. That gap — roughly 200 kW to 5 MW — is where most enterprises actually are, and it is the most under-served segment in AI infrastructure.
Here is how a 2 MW private cluster gets built in about a quarter, and where the time really goes.
Weeks 1–2: Sizing and Site
Start from the workload: training cadence, inference volume, model sizes, checkpoint frequency. That produces a GPU count, a fabric topology and a storage bandwidth target. In parallel, survey the site: available power at the switchgear, chilled or condenser water, slab rating, pathway for busway and pipework. The output is a design basis — one document that fixes the rack type, the count, the cooling method and the power topology — and a budgetary number.
Weeks 2–6: Engineering and Long-Lead Orders
Electrical one-lines, mechanical schematics, hydraulic design for the secondary loop, structural checks, the network design down to optic part numbers. Long-lead items go on order the moment the design basis is signed: GPUs and servers obviously, but also CDUs, busway and switchgear, which can have lead times as long as the compute. The schedule is set by whichever of those arrives last, so procurement is a critical-path activity, not a back-office one.
Weeks 6–12: Build
Busway and PDUs, pipework and CDUs, containment and cabling pathway go in while the racks are in transit. Racks arrive integrated where possible — liquid-ready, with nodes installed and cabled at the factory — and are placed, connected to the loop and the busway, and cabled to the fabric. Cabling is certified to 800G before the first power-on.
Weeks 12–14: Commission and Bring Up
Loop pressure test, flush and fill, flow balancing, thermal map under load. Electrical commissioning through to integrated systems testing. Then the cluster itself: burn-in, fabric benchmarks, storage tuning, scheduler, observability, and a reference training run at the agreed throughput. Handover is documentation and training, and the facility is in production.
Where It Slips
Three places, in our experience: a structural or power surprise found late because the survey was skipped; a long-lead item ordered after design rather than alongside it; and commissioning compressed to protect a go-live date. All three are avoidable, and all three are why the first two weeks are the most important of the fourteen.
AI Factory Design & Build
NVL72-class private clusters, from powered shell to first token.
