Nvidia's Next AI Chip Just Beat Blackwell by 2x in Public
Nvidia's Vera Rubin NVL72 debuted at MLPerf Inference v6.1, posting up to 2x the GB300's tokens per second before it's even shipping.
Nvidia usually controls when the world gets its first real look at a new chip's performance. This time, the world got there through a benchmark database. On September 16, 2026, MLCommons published the results of MLPerf Inference v6.1, and buried inside a 486-result dataset was the first public performance data for Vera Rubin NVL72, Nvidia's successor to the Blackwell GB300 platform. The chip isn't shipping to customers yet. It submitted under a "preview" label, which MLPerf allows specifically for systems in final pre-production. The numbers landed anyway.
The headline claims from Nvidia's own promotional materials say Vera Rubin NVL72 beats GB300 NVL72 by up to 3.7 times on Qwen3-VL and up to 2.5 times on DeepSeek-R1. Measured results from the actual submission, published by Nebius, which ran nine nodes of Vera Rubin against an eighteen-node GB300 reference, land closer to 1.6 times in offline mode and roughly 2.0 times in server mode. That's less than the marketing ceiling, and still a genuine, first-of-its-kind public benchmark for a chip nobody has bought yet.
What the Gap Between Claim and Measurement Actually Means
The difference between Nvidia's 2.5x claim and Nebius's 2.0x measured result comes down to hardware scale. Nebius submitted 36 Vera Rubin GPUs across nine nodes. The reference GB300 system used 72 GPUs across eighteen nodes, literally twice the silicon. Comparing those two configurations directly is what the benchmark reported. What Nvidia's own 2.5x claim assumes is a like-for-like rack comparison at equal node count, which hasn't been publicly submitted yet.
That gap isn't a scandal. Preview submissions routinely use whatever configuration a company has available rather than the full production rack, and MLPerf working-group chairs Miro Hodak and Frank Han published an explicit note in their September 17 analysis clarifying that any comparison between Rubin and older Nvidia generations, or against Google's TPU line, should be treated as speculation until equivalent-scale data exists. The headline claim Nvidia has made on its own benchmarks, up to 30 times higher throughput per megawatt on the SemiAnalysis AgentX agentic-coding benchmark, remains outside MLPerf's independent testing scope entirely, though the specific workload it targets, long-context KV-cache reuse at production agentic session scale, is real and measurable by anyone with access to the hardware.
The Architecture Behind the Numbers
Vera Rubin isn't an incremental upgrade to Blackwell. It's a seven-chip, five-rack architecture Nvidia designed from the ground up around what it calls the AI factory model: always-on inference at industrial scale, optimized for token revenue over the chip's useful life rather than raw peak throughput on a controlled benchmark. The individual Rubin GPU delivers 50 petaflops of FP4 performance and 22 terabytes per second of HBM4 memory bandwidth, the latter running at over 11 gigabits per second per pin, roughly 30% faster than AMD's equivalent HBM4 configurations. The full NVL72 rack houses 72 Rubin GPUs alongside 36 Vera CPUs, connected by NVLink 6, and a companion LPX rack integrates Groq 3 low-power inference accelerators for the decode-heavy, latency-sensitive part of serving trillion-parameter models.
The memory bandwidth figure matters more for how these chips are actually used in 2026 than the raw compute number does. Modern AI inference at production scale, including the long-context agentic sessions that have become standard in enterprise deployments, is almost entirely memory-bandwidth-bound rather than compute-bound. A chip that serves more tokens per second per watt on a 140,000-token context window is commercially more valuable than one that peaks higher on short synthetic prompts, which is exactly what the AgentX benchmark was designed to capture and what Nvidia claims Vera Rubin dramatically improves.
AMD and Crusoe Gave the Benchmark Its Other Headline
The same September 16 MLPerf round produced a separate, notable result from the other side of the competition. AMD submitted a 512-GPU Instinct MI355X cluster built and operated by Crusoe, the neocloud infrastructure company that tripled its valuation to $30 billion earlier this year. That cluster represents one of the largest single Instinct MI submissions in MLPerf history, proving the AMD architecture can actually be assembled at the scale that matters to hyperscale buyers rather than only in smaller research configurations. AMD also demonstrated software-only gains of 28 to 38 percent on MI355X hardware that wasn't changed between the last benchmark round and this one, a sign the company's inference software stack is still improving even without a new chip generation.
That combination, Vera Rubin's hardware debut and AMD's largest-ever cluster submission in the same round, made MLPerf v6.1 the most consequential AI inference benchmark cycle in several years. The 30 submitters and 120 systems representing 486 individual results also set participation records, a signal that inference optimization, rather than training, has become the competitive battleground the industry cares most about as agentic AI workloads drive a fundamental shift in what chip buyers actually need.
Why Memory Bandwidth Is Now the Only Number That Matters
One metric from OpenRouter's State of AI report, published alongside this MLPerf cycle, puts the architectural shift in sharp relief. Single agentic requests now consume 15 times the tokens of ordinary chat completions, a figure that has shifted what chip buyers optimize for far more decisively than any benchmark result on its own. A chip that excels at low-latency, short-context chat inference is increasingly misaligned with what production AI systems actually run. Vera Rubin's specific design choices, the HBM4 bandwidth spec, the KV-cache reuse optimization, the Groq 3 LPU companion for decode, all reflect Nvidia's own read of where production workloads have moved.
SK Hynix's entire $720 billion memory capacity expansion is predicated on the same shift. Chips that can serve longer contexts faster, at lower power cost, need more bandwidth per GPU, which requires more HBM4 per chip, which requires more HBM wafer production than the entire global industry currently has capacity to deliver. Vera Rubin's 22 terabytes per second bandwidth spec doesn't exist in isolation from the memory supply chain story โ it's part of why the demand for HBM4 that SK Hynix is building toward 2028 is already sold out before the fabs are built.
The Preview Label Does Real Work Here
MLPerf's preview category carries a specific, meaningful guarantee: any system that submits in preview must achieve full certified compliance by the following benchmark round if it wants to keep appearing in the results. That's not a soft expectation. It's a public commitment Nvidia made by submitting, and it constrains the chip's commercial timeline more precisely than any press release would. Vera Rubin NVL72 will need to be in certified, reproducible production by the next MLPerf cycle to maintain its benchmark standing. Given the scale of Vera Rubin orders already on Nvidia's books, including Jensen Huang's confirmation of $1 trillion in combined Blackwell and Vera Rubin purchase orders through 2027, that timeline appears to be tracking exactly as planned.
What the MLPerf debut actually confirmed, beyond the specific numbers, is that Vera Rubin performs meaningfully better than Blackwell in at least the configurations tested so far, is real hardware that can be benchmarked rather than a slide deck, and is on a timeline credible enough that the company was willing to stake its public benchmark standing on completing full certification within the next cycle. For an industry increasingly focused on inference cost rather than training peak, that combination matters considerably more than any individual performance ratio the promotional materials chose to headline.
Written by
Mr. Aayush Bhatt
Software Engineer with in depth understanding of buliding softwares and Tech.




