NVIDIA's Blackwell GB300 and GB200 graphics processors continue to demonstrate leading performance in various artificial intelligence workloads through continuous technical optimization. As the new generation of AI platforms continues to be deployed around the world, the outside world may have speculated that NVIDIA will shift its focus to the Rubin architecture. However, it turns out that the Blackwell series, which has been put into practice in countless data centers, like the Hopper architecture of the year, is still releasing amazing potential through continuous software upgrades. Among them, the GB300 continues to set new performance highs in artificial intelligence computing tasks.

According to official data released by NVIDIA, the continuous optimization of the Blackwell platform has greatly improved the overall performance and energy efficiency ratio. In just 3 months, NVIDIA increased throughput per megawatt by as much as 4x on the same GB200 NVL72 configuration running DeepSeek R1 0528 - 1K/1K workloads. During this time, the company added 38 major optimizations to the Blackwell platform by running more than 250,000 simulation configurations and consuming a total of 1.4 million GPU test hours. It is worth mentioning that more than 90% of the optimization results can be applied to other artificial intelligence models, which fully demonstrates NVIDIA's commitment to maintaining deep investment in existing AI platforms as new platforms gradually become more popular.

On the other hand, NVIDIA's GB300 "Blackwell Ultra" platform continues to dominate the field of artificial intelligence training. When using 256 GPUs to run the DeepSeek-V3 671B model, its Megatron Core core set an astonishing record of 1,648 Terops per GPU, achieving a 3x performance leap compared to GB200 606 Terops. Using GB300 NVL72 to pre-train DeepSeek-V3 671B can achieve a world-record operating efficiency of 1648 Terops per GPU, allowing the same training task to achieve a given performance with only a fraction of the hardware scale. At the same time, the performance of the same GB300 NVL72 system has also increased by up to 1.5 times in the past 6 months after continuous optimization.

In addition, NVIDIA has also launched in-depth cooperation with PyTorch and the JAX open source community to promote the full implementation of various optimization results in its artificial intelligence product matrix. When running the DeepSeek-V3 671B model using PyTorch's native training stack TorchTitan, the Blackwell Ultra rack achieved up to 6x performance improvement without any initial optimization. The JAX framework has achieved even more significant growth. Its speed has soared to 1025 Terops per GPU, and the throughput of a single GPU can reach up to 4082 Tokens per second. Compared with the baseline in January this year, it has achieved a performance leap of 10 times.

2026-07-22_7-21-04-1.jpg2026-07-22_7-20-55.jpg2026-07-22_7-20-47.jpg2026-07-22_7-20-37.jpgimage-57.webp

Not only that, NVIDIA has also achieved world-class scaling performance from 256 to 1024 GPUs in all mainstream frameworks for pre-training. When reaching the scale of 1024 GPUs, Megatron Core can maintain an expansion efficiency of up to 98.5%, while the expansion efficiency of TorchTitan and the JAX framework also maintains a level of 97%. The core component that supports this excellent expansion performance is the NVIDIA 800Gb/s scale-out network chip integrated inside each NVL72 rack. From model training to inference deployment, NVIDIA is constantly setting new benchmarks in the field of artificial intelligence. With the Rubin platform already on the market and bringing greater performance gains, other competitors are undoubtedly facing considerable pressure to catch up.