Nvidia Vera Rubin NVL72 Shows Massive Inference Performance and TCO Gains Over Blackwell
SemiAnalysis released an architectural and total cost of ownership analysis comparing Nvidia's Vera Rubin NVL72 against the GB200 NVL72. Early engineering data from CoreWeave indicates that the Vera Rubin NVL72 running DeepSeek R1 delivers 5.4 times the performance per megawatt and 5 times the performance per dollar compared to current GB200 NVL72 systems. Additionally, Nvidia has launched its first public Rubin software stack with CUDA 13.4, upstreaming pull requests to PyTorch, vLLM, and OpenAI Triton to allow the reuse of Blackwell kernels. Rubin also benefits from a simpler, cableless compute tray design to accelerate its production ramp. (source: https://newsletter.semianalysis.com/p/vera-rubin-nvl72-vs-gb200-nvl72-inference)