Vera Rubin NVL72: What next gen training infra means
By Win.AI Editorial

Vera Rubin NVL72 is NVIDIA’s answer to the limits of server-scale GPUs: a purpose-built, rack-scale machine that hides cable complexity, packs 72 Rubin GPUs and 36 Vera CPUs into a single liquid-cooled chassis, and moves bandwidth inside the rack so larger models run faster per watt. That design changes the arithmetic of training economics while forcing a sharper choice between cloud, colocation and buying your own hardware.
Vera Rubin NVL72 hardware and network innovations
NVIDIA’s product pages and reference architecture documentation enumerate the concrete changes. A GB300 NVL72 rack contains 72 Rubin GPUs and 36 Grace/Vera CPUs, a NVLink fabric rated at roughly 260 terabytes per second across the rack which works out to about 3.6 terabytes per second of all-to-all bandwidth per GPU, and hot-swappable NVLink switch trays that remove long external cables. NVIDIA claims up to 10 times more tokens per megawatt compared with the previous GB200 generation. The rack is fully liquid cooled and commonly quoted power figures run in the 120 to 155 kilowatt range depending on configuration and workload, so facility power and cooling are first-order constraints. Sources: NVIDIA NVL72 product and developer pages, HPE and Lenovo GB300 documentation.
Practical trade-offs for buyers
If you operate at hyperscaler scale the math is simple. Higher efficiency per megawatt reduces energy spend and reduces wall time for multi-trillion parameter training runs. The NVLink fabric reduces inter-GPU synchronization overhead, which lowers training wall clock time for tightly coupled model parallel work. If you are a cloud, that translates into more throughput per data center bay.
For enterprises and labs the trade-offs complicate acquisition decisions. The hardware demands bespoke power distribution, 50 volt DC busbars or multiple 33 kW power shelves, chilled water or closed-loop CDUs, and staff experienced in liquid-cooled rack service. Those enable performance but add fixed facility upgrades that are hard to amortize under small or bursty usage.
One obvious counterargument is vendor diversification. Alternative accelerators and systems such as purpose-built wafers or Cerebras-class appliances have different integration points and can avoid NVLink dependency. Large cloud providers have already mixed accelerator types in production, which reduces single-vendor risk. See the historical example of large cloud deployments of alternative accelerators for context Amazon deploys Cerebras technology 25 times faster than.
Decision guide and lessons observed
If you need sustained, predictable large-scale training capacity and operate hundreds of racks the NVL72 path will usually lower run time and energy cost per parameter. If your needs are intermittent or constrained to dozens of racks rent first. Colocation providers will offer midpoints but expect premium pricing for sufficient power and CDU capacity.
We observed three recurring operational patterns while studying rack-scale NVL systems. First, power design becomes the project, not the rack. Second, software tooling that exploits rack-level NVLink is the gating factor for real throughput gains. Third, installation and servicing speed matters because a single failed tray can stop a multi-rack job.
NVIDIA’s NVL72 changes the center of gravity for training infrastructure. It favors organizations that can invest in campus or hyperscale power and cooling and that plan software to use tight all-to-all links. For everyone else the sensible path is staged adoption: cloud or colocation for initial scale and a targeted on-prem commitment only once utilization stays high and predictable.




