Nvidia said its Vera Rubin platform, carrying 288 gigabytes of HBM4 per GPU with 22 terabytes per second of memory bandwidth and 336 billion transistors, moved into full production for a second-half 2026 ramp.1 Thunder Compute 2026-03-05 Each Rubin GPU carries 288 GB of HBM4 at 22 TB/s and 336 billion transistors; the platform entered full production at CES 2026 with cloud availability in H2 2026, and the 10x inference cost claim applies to mixture-of-experts models at long sequence lengths, with dense-model gains nearer 2 to 3x. Open source The company's headline claim is that Vera Rubin delivers up to 5x greater inference performance and 10x lower cost per token than the current Blackwell generation.3 Tom's Hardware 2026-01-06 Nvidia promised the Vera Rubin NVL72 will deliver up to 5x greater inference performance and 10x lower cost per token than Blackwell, coming in the second half of 2026. Open source The stake is which metric now governs the AI chip market. We assess, with moderate confidence, that Rubin's significance is its inference economics rather than peak training throughput, and that the memory subsystem, not raw compute, is what does the work behind the cost claim.

The number that matters is memory bandwidth

The load-bearing specification is not the transistor count, it is the memory. Each Rubin GPU triples per-GPU memory bandwidth to 22 terabytes per second using HBM4, a 2.8x improvement over Blackwell's roughly 8 terabytes per second on HBM3e.2 Barrack AI 2026-03-06 Rubin offers 288 GB of HBM4 per GPU at 22 TB/s, a 2.8x bandwidth improvement over Blackwell 8 TB/s, 336 billion transistors and 50 petaFLOPS of FP4 inference per chip, at 1,800 to 2,300W requiring liquid cooling. Open source Inference at scale is bound by how fast a chip can move weights and context to the compute units, so bandwidth and capacity, not FLOPS, are what cut cost per token. The 288 gigabytes of HBM4 per GPU let a chip hold larger models and longer contexts in fast memory, and Nvidia pairs that with a doubling of rack-scale interconnect to 260 terabytes per second in the NVL72 configuration.1 Thunder Compute 2026-03-05 Each Rubin GPU carries 288 GB of HBM4 at 22 TB/s and 336 billion transistors; the platform entered full production at CES 2026 with cloud availability in H2 2026, and the 10x inference cost claim applies to mixture-of-experts models at long sequence lengths, with dense-model gains nearer 2 to 3x. Open source The chip still posts large compute figures, 50 petaFLOPS of FP4 inference per GPU, but the framing across the platform is inference, and it comes at a steep 1,800 to 2,300 watt draw that requires liquid cooling.2 Barrack AI 2026-03-06 Rubin offers 288 GB of HBM4 per GPU at 22 TB/s, a 2.8x bandwidth improvement over Blackwell 8 TB/s, 336 billion transistors and 50 petaFLOPS of FP4 inference per chip, at 1,800 to 2,300W requiring liquid cooling. Open source

That framing is itself the signal. Nvidia is marketing Rubin on cost per token and inference performance, not on training records, which tracks where the money in AI is moving: from training a model once to serving it billions of times.

Read the fine print on 10x

The 10x cost claim deserves the scrutiny an intelligence reader owes a vendor number. The headline reduction in inference token cost is not a blanket result: it applies specifically to mixture-of-experts models at long sequence lengths, the workloads that benefit most from the bandwidth and interconnect gains.1 Thunder Compute 2026-03-05 Each Rubin GPU carries 288 GB of HBM4 at 22 TB/s and 336 billion transistors; the platform entered full production at CES 2026 with cloud availability in H2 2026, and the 10x inference cost claim applies to mixture-of-experts models at long sequence lengths, with dense-model gains nearer 2 to 3x. Open source For dense-model inference, the realistic improvement is closer to 2 to 3x rather than 10x.1 Thunder Compute 2026-03-05 Each Rubin GPU carries 288 GB of HBM4 at 22 TB/s and 336 billion transistors; the platform entered full production at CES 2026 with cloud availability in H2 2026, and the 10x inference cost claim applies to mixture-of-experts models at long sequence lengths, with dense-model gains nearer 2 to 3x. Open source This matters because it means the marketing figure describes a best case tied to a specific architecture and context regime, and buyers running dense models or short contexts should plan around the lower number. We assess, with moderate confidence, that the real generational gain most customers see will land between 2x and 10x depending on workload, and that the 10x is honest only with its conditions attached.

Who gains and who loses

Nvidia gains the most: by defining the competitive axis as inference cost per token and shipping a memory-heavy part built for it, it extends its position into the phase of the market that is growing fastest, and does so before hyperscaler custom silicon closes the gap.3 Tom's Hardware 2026-01-06 Nvidia promised the Vera Rubin NVL72 will deliver up to 5x greater inference performance and 10x lower cost per token than Blackwell, coming in the second half of 2026. Open source The HBM4 suppliers gain directly, since a platform demanding 288 gigabytes per GPU pulls enormous high-bandwidth memory volume, and HBM is the scarce input; Nvidia has qualified multiple memory makers to feed the ramp.1 Thunder Compute 2026-03-05 Each Rubin GPU carries 288 GB of HBM4 at 22 TB/s and 336 billion transistors; the platform entered full production at CES 2026 with cloud availability in H2 2026, and the 10x inference cost claim applies to mixture-of-experts models at long sequence lengths, with dense-model gains nearer 2 to 3x. Open source

The pressured parties are the custom-silicon programs at Microsoft, Amazon, and Google that pitch themselves on better inference cost per token, exactly the ground Rubin now claims with a large bandwidth advantage.2 Barrack AI 2026-03-06 Rubin offers 288 GB of HBM4 per GPU at 22 TB/s, a 2.8x bandwidth improvement over Blackwell 8 TB/s, 336 billion transistors and 50 petaFLOPS of FP4 inference per chip, at 1,800 to 2,300W requiring liquid cooling. Open source Their in-house chips must now beat a moving target. Data center operators also absorb a cost: at 1,800 to 2,300 watts per GPU with mandatory liquid cooling, Rubin raises the power-density and cooling bar, which favors newer, better-provisioned facilities and penalizes older ones.2 Barrack AI 2026-03-06 Rubin offers 288 GB of HBM4 per GPU at 22 TB/s, a 2.8x bandwidth improvement over Blackwell 8 TB/s, 336 billion transistors and 50 petaFLOPS of FP4 inference per chip, at 1,800 to 2,300W requiring liquid cooling. Open source

The counter-case

The thesis that Rubin locks in Nvidia for the inference era could fail on two fronts. First, the cost advantage is workload-specific: if the bulk of real inference demand is dense models or shorter contexts, the effective gain is the 2 to 3x figure, not 10x, which narrows the lead custom silicon has to overcome.1 Thunder Compute 2026-03-05 Each Rubin GPU carries 288 GB of HBM4 at 22 TB/s and 336 billion transistors; the platform entered full production at CES 2026 with cloud availability in H2 2026, and the 10x inference cost claim applies to mixture-of-experts models at long sequence lengths, with dense-model gains nearer 2 to 3x. Open source Second, the power and cooling requirements are demanding enough that deployment could lag the production ramp, since not every data center can host 2,300 watt liquid-cooled GPUs at rack scale.2 Barrack AI 2026-03-06 Rubin offers 288 GB of HBM4 per GPU at 22 TB/s, a 2.8x bandwidth improvement over Blackwell 8 TB/s, 336 billion transistors and 50 petaFLOPS of FP4 inference per chip, at 1,800 to 2,300W requiring liquid cooling. Open source For the bull thesis to break, a competitor's inference chip would need to match Rubin's cost per token on common workloads while drawing less power, or the deployment friction would need to slow the H2 2026 ramp materially. Both are plausible, neither is yet evident.

What to watch

  • Volume shipments hit the H2 2026 window. If Rubin reaches broad cloud availability across major providers in the second half of 2026 as stated, the ramp is on schedule; a slip into 2027 would ease the pressure on rival silicon.3 Tom's Hardware 2026-01-06 Nvidia promised the Vera Rubin NVL72 will deliver up to 5x greater inference performance and 10x lower cost per token than Blackwell, coming in the second half of 2026. Open source
  • Independent cost-per-token tests land between 2x and 10x. Third-party inference benchmarks within a couple of quarters would show where the real gain falls; results clustering near 2 to 3x would confirm the 10x is a narrow best case.1 Thunder Compute 2026-03-05 Each Rubin GPU carries 288 GB of HBM4 at 22 TB/s and 336 billion transistors; the platform entered full production at CES 2026 with cloud availability in H2 2026, and the 10x inference cost claim applies to mixture-of-experts models at long sequence lengths, with dense-model gains nearer 2 to 3x. Open source
  • HBM4 supply keeps pace. If qualified memory makers deliver enough HBM4 to feed a 288 gigabyte-per-GPU platform at volume, the ramp is unconstrained; an HBM shortage would cap it regardless of demand.1 Thunder Compute 2026-03-05 Each Rubin GPU carries 288 GB of HBM4 at 22 TB/s and 336 billion transistors; the platform entered full production at CES 2026 with cloud availability in H2 2026, and the 10x inference cost claim applies to mixture-of-experts models at long sequence lengths, with dense-model gains nearer 2 to 3x. Open source
  • Cooling and power gate deployment. If operators report difficulty hosting 1,800 to 2,300 watt liquid-cooled parts, adoption lags production.2 Barrack AI 2026-03-06 Rubin offers 288 GB of HBM4 per GPU at 22 TB/s, a 2.8x bandwidth improvement over Blackwell 8 TB/s, 336 billion transistors and 50 petaFLOPS of FP4 inference per chip, at 1,800 to 2,300W requiring liquid cooling. Open source The metric that decides Nvidia's hold on the inference era is not peak FLOPS, it is verified cost per token in customers' own workloads.