This is NVIDIA's Data Center Deep Learning Product Performance Hub — a centralized resource for reproducible AI performance benchmarks across NVIDIA's latest data center GPUs.
The days of raw speed being the only metric that matters are behind us. Now it’s about throughput, efficiency, and economics at scale. As AI evolves from providing one-shot answers to engaging in multi-step reasoning, the demand for inference and its underlying economics is increasing.. This shift significantly boosts compute demand due to the generation of far more tokens per query. Metrics such as tokens per watt, cost per million tokens, and tokens per second per user are crucial alongside throughput.
For power-limited AI factories, NVIDIA's continuous software improvements translate into higher token revenue over time, underscoring the importance of our technological advancements.
Pareto curves illustrate how NVIDIA Vera Rubin and provides the best balance across the full spectrum of production priorities, including:
Cost
Energy efficiency
Throughput
Responsiveness
Optimizing systems for a single scenario can limit deployment flexibility, leading to inefficiencies at other points on the curve. NVIDIA’s full-stack design approach ensures efficiency and value across multiple real-life production scenarios. Blackwell’s leadership stems from its extreme hardware-software co-design, embodying a full-stack architecture built for speed, efficiency, and scalability.
Explore the methodology used to obtain these results and learn how to replicate the tests by executing Benchmarking Recipes yourself.
| Network | Throughput | GPU | Server | GPU Version | QSL Size | Target Accuracy | Dataset |
|---|---|---|---|---|---|---|---|
| DeepSeek R1 | 1,183,327 tokens/sec | 72x VR200† | NVIDIA Vera Rubin NVL72 (72x VR200-288GB_aarch64) | NVIDIA Vera Rubin NVL72 | 4388 | 99% of FP16 (exact match 81.9132%) | mlperf_deepseek_r1 |
| 2,705,130 tokens/sec | 288x GB300 | NVIDIA GB300 NVL72 (72x GB300-288GB_aarch64, TensorRT) | NVIDIA GB300 NVL72 | 4388 | 99% of FP16 (exact match 81.9132%) | mlperf_deepseek_r1 | |
| 2,091,190 tokens/sec | 288x GB200 | Azure GB200 NVL72 (288x GB200-186GB_aarch64, TensorRT) | NVIDIA GB200 NVL72 | 4388 | 99% of FP16 (exact match 81.9132%) | mlperf_deepseek_r1 | |
| 137,067 tokens/sec | 16x B300 | Cisco_UCS_B300x16_G200_4x2_RO | NVIDIA B300 | 4388 | 99% of FP16 (exact match 81.9132%) | mlperf_deepseek_r1 | |
| 58,913 tokens/sec | 8x B200 | CoreWeave B200 SXM (8x B200-SXM-180GB) | NVIDIA B200 | 4388 | 99% of FP16 (exact match 81.9132%) | mlperf_deepseek_r1 | |
| gpt-oss 120B | 1,197,710 tokens/sec | 72x GB300 | CoreWeave GB300 NVL72 (72x GB300-288GB_aarch64) | NVIDIA GB300 NVL72 | 6396 | 99% of 83.13% | AIME25, GPQA Diamond, LiveCodeBench v6 |
| 911,566 tokens/sec | 72x GB200 | CoreWeave GB200 NVL72 (72x GB200-186GB_aarch64) | NVIDIA GB200 NVL72 | 6396 | 99% of 83.13% | AIME25, GPQA Diamond, LiveCodeBench v6 | |
| 220,527 tokens/sec | 16x B300 | Cisco_UCS_B300x16_G200_4x2_RO | NVIDIA B300 | 6396 | 99% of 83.13% | AIME25, GPQA Diamond, LiveCodeBench v6 | |
| 91,487 tokens/sec | 8x B200 | CoreWeave B200 SXM (8x B200-SXM-180GB) | NVIDIA B200 | 6396 | 99% of 83.13% | AIME25, GPQA Diamond, LiveCodeBench v6 | |
| Qwen3-VL 235B | 2,393 samples/sec | 72x VR200† | NVIDIA Vera Rubin NVL72 (72x VR200-288GB_aarch64) | NVIDIA Vera Rubin NVL72 | 48289 | 99% of BF16 (Category Hierarchical F1 Score >= 0.7824) | Shopify Product Catalogue |
| 1,305 samples/sec | 72x GB300 | NVIDIA GB300 NVL72 (72x GB300-288GB_aarch64, TensorRT) | NVIDIA GB300 NVL72 | 48289 | 99% of BF16 (Category Hierarchical F1 Score >= 0.7824) | Shopify Product Catalogue | |
| 230 samples/sec | 16x GB200 | NVIDIA GB200 NVL72 (16x GB200-186GB_aarch64, Dynamo) | NVIDIA GB200 NVL72 | 48289 | 99% of BF16 (Category Hierarchical F1 Score >= 0.7824) | Shopify Product Catalogue | |
| 132 samples/sec | 8x B300 | G894-ZD3-AAX7 | NVIDIA B300 | 48289 | 99% of BF16 (Category Hierarchical F1 Score >= 0.7824) | Shopify Product Catalogue | |
| 102 samples/sec | 8x B200 | Lambda Cloud 8x B200-SXM-180GB | NVIDIA B200 | 48289 | 99% of BF16 (Category Hierarchical F1 Score >= 0.7824) | Shopify Product Catalogue | |
| Llama2 70B | 1,136,100 tokens/sec | 72x GB300 | CoreWeave GB300 NVL72 (72x GB300-288GB_aarch64) | NVIDIA GB300 NVL72 | 24576 | 99.9% of FP32 (rouge1=44.4312, rouge2=22.0352, rougeL=28.6162). Additionally, the generation length of the tokens per sample should be more than 90% of the reference (tokens_per_sample=294.45) | OpenOrca (max_seq_len=1024) |
| 894,002 tokens/sec | 72x GB200 | CoreWeave GB200 NVL72 (72x GB200-186GB_aarch64) | NVIDIA GB200 NVL72 | 24576 | 99.9% of FP32 (rouge1=44.4312, rouge2=22.0352, rougeL=28.6162). Additionally, the generation length of the tokens per sample should be more than 90% of the reference (tokens_per_sample=294.45) | OpenOrca (max_seq_len=1024) | |
| 227,570 tokens/sec | 16x B300 | Cisco_UCS_B300x16_G200_4x2_RO | NVIDIA B300 | 24576 | 99.9% of FP32 (rouge1=44.4312, rouge2=22.0352, rougeL=28.6162). Additionally, the generation length of the tokens per sample should be more than 90% of the reference (tokens_per_sample=294.45) | OpenOrca (max_seq_len=1024) | |
| 102,703 tokens/sec | 8x B200 | CoreWeave B200 SXM (8x B200-SXM-180GB) | NVIDIA B200 | 24576 | 99.9% of FP32 (rouge1=44.4312, rouge2=22.0352, rougeL=28.6162). Additionally, the generation length of the tokens per sample should be more than 90% of the reference (tokens_per_sample=294.45) | OpenOrca (max_seq_len=1024) | |
| Llama3.1 8B | 171,114 tokens/sec | 8x B300 | Cisco UCS C880A M8 (8x NVIDIA B300-SXM-270GB, TensorRT) | NVIDIA B300 | 13368 | 99% of FP32 and 99.9% of FP32 (rouge1=42.9865, rouge2=20.1235, rougeL=29.9881). Additionally, for both cases the total generation length of the texts should be more than 90% of the reference (gen_len=8167644) | CNN Dailymail (v3.0.0, max_seq_len=2048) |
| Whisper | 53,267 samples/sec | 8x B300 | NVIDIA DGX B300 (8x B300-SXM-270GB, TensorRT) | NVIDIA B300 | 1633 | 99% of FP32 and 99.9% of FP32 (WER=2.0671%) | LibriSpeech |
| Wan2.2 | 0.655 samples/sec | 72x GB300 | NVIDIA GB300 NVL72 (72x GB300-288GB_aarch64, TensorRT) | NVIDIA GB300 NVL72 | 248 | 99% of BF16 (VBench score >= 69.7752) | VBench prompts |
| 0.496 samples/sec | 72x GB200 | NVIDIA GB200 NVL72 (72x GB200-186GB_aarch64, TensorRT) | NVIDIA GB200 NVL72 | 248 | 99% of BF16 (VBench score >= 69.7752) | VBench prompts | |
| 0.078 samples/sec | 8x B300 | NVIDIA DGX B300 (8x B300-SXM-270GB, TensorRT) | NVIDIA B300 | 248 | 99% of BF16 (VBench score >= 69.7752) | VBench prompts |
| Network | Throughput | GPU | Server | GPU Version | QSL Size | Target Accuracy | MLPerf Server Latency
Constraints (ms) |
Dataset |
|---|---|---|---|---|---|---|---|---|
| DeepSeek R1 | 1,175,890 tokens/sec | 72x VR200† | NVIDIA Vera Rubin NVL72 (72x VR200-288GB_aarch64) | NVIDIA Vera Rubin NVL72 | 4388 | 99% of FP16 (exact match 81.9132%) | TTFT/TPOT: 2000 ms/80 ms | mlperf_deepseek_r1 |
| 2,028,030 tokens/sec | 288x GB300 | NVIDIA GB300 NVL72 (72x GB300-288GB_aarch64, TensorRT) | NVIDIA GB300 NVL72 | 4388 | 99% of FP16 (exact match 81.9132%) | TTFT/TPOT: 2000 ms/80 ms | mlperf_deepseek_r1 | |
| 1,598,510 tokens/sec | 288x GB200 | Azure GB200 NVL72 (288x GB200-186GB_aarch64, TensorRT) | NVIDIA GB200 NVL72 | 4388 | 99% of FP16 (exact match 81.9132%) | TTFT/TPOT: 2000 ms/80 ms | mlperf_deepseek_r1 | |
| 88,916 tokens/sec | 16x B300 | Cisco_UCS_B300x16_G200_4x2_RO | NVIDIA B300 | 4388 | 99% of FP16 (exact match 81.9132%) | TTFT/TPOT: 2000 ms/80 ms | mlperf_deepseek_r1 | |
| 56,121 tokens/sec | 8x B200 | CoreWeave B200 SXM (8x B200-SXM-180GB) | NVIDIA B200 | 4388 | 99% of FP16 (exact match 81.9132%) | TTFT/TPOT: 2000 ms/80 ms | mlperf_deepseek_r1 | |
| gpt-oss 120B | 1,160,490 tokens/sec | 72x GB300 | CoreWeave GB300 NVL72 (72x GB300-288GB_aarch64) | NVIDIA GB300 NVL72 | 6396 | 99% of 83.13% | TTFT/TPOT: 3000 ms/80 ms | AIME25, GPQA Diamond, LiveCodeBench v6 |
| 901,058 tokens/sec | 72x GB200 | CoreWeave GB200 NVL72 (72x GB200-186GB_aarch64) | NVIDIA GB200 NVL72 | 6396 | 99% of 83.13% | TTFT/TPOT: 3000 ms/80 ms | AIME25, GPQA Diamond, LiveCodeBench v6 | |
| 214,774 tokens/sec | 16x B300 | Cisco_UCS_B300x16_G200_4x2_RO | NVIDIA B300 | 6396 | 99% of 83.13% | TTFT/TPOT: 3000 ms/80 ms | AIME25, GPQA Diamond, LiveCodeBench v6 | |
| 90,205 tokens/sec | 8x B200 | CoreWeave B200 SXM (8x B200-SXM-180GB) | NVIDIA B200 | 6396 | 99% of 83.13% | TTFT/TPOT: 3000 ms/80 ms | AIME25, GPQA Diamond, LiveCodeBench v6 | |
| Qwen3-VL 235B | 2,323 queries/sec | 72x VR200† | NVIDIA Vera Rubin NVL72 (72x VR200-288GB_aarch64) | NVIDIA Vera Rubin NVL72 | 48289 | 99% of BF16 (Category Hierarchical F1 Score >= 0.7824) | 12 s | Shopify Product Catalogue |
| 1,210 queries/sec | 72x GB300 | NVIDIA GB300 NVL72 (72x GB300-288GB_aarch64, TensorRT) | NVIDIA GB300 NVL72 | 48289 | 99% of BF16 (Category Hierarchical F1 Score >= 0.7824) | 12 s | Shopify Product Catalogue | |
| 199 queries/sec | 16x GB200 | NVIDIA GB200 NVL72 (16x GB200-186GB_aarch64, Dynamo) | NVIDIA GB200 NVL72 | 48289 | 99% of BF16 (Category Hierarchical F1 Score >= 0.7824) | 12 s | Shopify Product Catalogue | |
| 116 queries/sec | 8x B300 | G894-ZD3-AAX7 | NVIDIA B300 | 48289 | 99% of BF16 (Category Hierarchical F1 Score >= 0.7824) | 12 s | Shopify Product Catalogue | |
| 69.415 queries/sec | 8x B200 | Lambda Cloud 8x B200-SXM-180GB | NVIDIA B200 | 48289 | 99% of BF16 (Category Hierarchical F1 Score >= 0.7824) | 12 s | Shopify Product Catalogue | |
| Llama2 70B | 944,902 tokens/sec | 72x GB300 | CoreWeave GB300 NVL72 (72x GB300-288GB_aarch64) | NVIDIA GB300 NVL72 | 24576 | 99.9% of FP32 (rouge1=44.4312, rouge2=22.0352, rougeL=28.6162). Additionally, the generation length of the tokens per sample should be more than 90% of the reference (tokens_per_sample=294.45) | TTFT/TPOT: 2000 ms/200 ms | OpenOrca (max_seq_len=1024) |
| 910,621 tokens/sec | 72x GB200 | CoreWeave GB200 NVL72 (72x GB200-186GB_aarch64) | NVIDIA GB200 NVL72 | 24576 | 99.9% of FP32 (rouge1=44.4312, rouge2=22.0352, rougeL=28.6162). Additionally, the generation length of the tokens per sample should be more than 90% of the reference (tokens_per_sample=294.45) | TTFT/TPOT: 2000 ms/200 ms | OpenOrca (max_seq_len=1024) | |
| 179,597 tokens/sec | 16x B300 | Cisco_UCS_B300x16_G200_4x2_RO | NVIDIA B300 | 24576 | 99.9% of FP32 (rouge1=44.4312, rouge2=22.0352, rougeL=28.6162). Additionally, the generation length of the tokens per sample should be more than 90% of the reference (tokens_per_sample=294.45) | TTFT/TPOT: 2000 ms/200 ms | OpenOrca (max_seq_len=1024) | |
| 102,398 tokens/sec | 8x B200 | CoreWeave B200 SXM (8x B200-SXM-180GB) | NVIDIA B200 | 24576 | 99.9% of FP32 (rouge1=44.4312, rouge2=22.0352, rougeL=28.6162). Additionally, the generation length of the tokens per sample should be more than 90% of the reference (tokens_per_sample=294.45) | TTFT/TPOT: 2000 ms/200 ms | OpenOrca (max_seq_len=1024) | |
| Llama3.1 8B | 173,950 tokens/sec | 8x B300 | Cisco UCS C880A M8 (8x NVIDIA B300-SXM-270GB, TensorRT) | NVIDIA B300 | 13368 | 99% of FP32 and 99.9% of FP32 (rouge1=42.9865, rouge2=20.1235, rougeL=29.9881). Additionally, for both cases the total generation length of the texts should be more than 90% of the reference (gen_len=8167644) | TTFT/TPOT: 2000 ms/100 ms | CNN Dailymail (v3.0.0, max_seq_len=2048) |
| Wan2.2** | 5.73 seconds | 72x GB200 | NVIDIA GB200 NVL72 (72x GB200-186GB_aarch64, TensorRT) | NVIDIA GB200 NVL72 | 248 | 99% of BF16 (VBench score >= 69.7752) | N/A | VBench prompts |
| 5.81 seconds | 72x GB300 | NVIDIA GB300 NVL72 (72x GB300-288GB_aarch64, TensorRT) | NVIDIA GB300 NVL72 | 248 | 99% of BF16 (VBench score >= 69.7752) | N/A | VBench prompts | |
| 17.98 seconds | 8x B300 | NVIDIA DGX B300 (8x B300-SXM-270GB, TensorRT) | NVIDIA B300 | 248 | 99% of BF16 (VBench score >= 69.7752) | N/A | VBench prompts |
| Network | Throughput | GPU | Server | GPU Version | QSL Size | Target Accuracy | MLPerf Server Latency
Constraints (ms) |
Dataset |
|---|---|---|---|---|---|---|---|---|
| DeepSeek R1 | 652,750 tokens/sec | 72x VR200† | NVIDIA Vera Rubin NVL72 (72x VR200-288GB_aarch64) | NVIDIA Vera Rubin NVL72 | 4388 | 99% of FP16 (exact match 81.9132%) | TTFT/TPOT: 1500 ms/15 ms | mlperf_deepseek_r1 |
| 260,098 tokens/sec | 72x GB300 | Azure GB300 NVL72 (72x GB300-288GB_aarch64 TensorRT) | NVIDIA GB300 NVL72 | 4388 | 99% of FP16 (exact match 81.9132%) | TTFT/TPOT: 1500 ms/15 ms | mlperf_deepseek_r1 | |
| 240,274 tokens/sec | 72x GB200 | NVIDIA GB200 NVL72 (72x GB200-186GB_aarch64, TensorRT) | NVIDIA GB200 NVL72 | 4388 | 99% of FP16 (exact match 81.9132%) | TTFT/TPOT: 1500 ms/15 ms | mlperf_deepseek_r1 | |
| gpt-oss 120B | 669,305 tokens/sec | 72x GB300 | NVIDIA GB300 NVL72 (72x GB300-288GB_aarch64, TensorRT) | NVIDIA GB300 NVL72 | 6396 | 99% of 83.13% | TTFT/TPOT: 2000 ms/20 ms | AIME25, GPQA Diamond, LiveCodeBench v6 |
| 625,205 tokens/sec | 72x GB200 | NVIDIA GB200 NVL72 (72x GB200-186GB_aarch64, TensorRT) | NVIDIA GB200 NVL72 | 6396 | 99% of 83.13% | TTFT/TPOT: 2000 ms/20 ms | AIME25, GPQA Diamond, LiveCodeBench v6 | |
| Qwen3-VL 235B | 1,307 queries/sec | 72x VR200† | NVIDIA Vera Rubin NVL72 (72x VR200-288GB_aarch64) | NVIDIA Vera Rubin NVL72 | 48289 | 99% of BF16 (Category Hierarchical F1 Score >= 0.7824) | 1.5 s | Shopify Product Catalogue |
| 349 queries/sec | 72x GB300 | NVIDIA GB300 NVL72 (72x GB300-288GB_aarch64, TensorRT) | NVIDIA GB300 NVL72 | 48289 | 99% of BF16 (Category Hierarchical F1 Score >= 0.7824) | 1.5 s | Shopify Product Catalogue | |
| 39.851 queries/sec | 16x GB200 | NVIDIA GB200 NVL72 (16x GB200-186GB_aarch64, Dynamo) | NVIDIA GB200 NVL72 | 48289 | 99% of BF16 (Category Hierarchical F1 Score >= 0.7824) | 1.5 s | Shopify Product Catalogue | |
| 12.882 queries/sec | 8x B300 | G894-ZD3-AAX7 | NVIDIA B300 | 48289 | 99% of BF16 (Category Hierarchical F1 Score >= 0.7824) | 1.5 s | Shopify Product Catalogue | |
| Llama3.1 8B | 129,168 tokens/sec | 8x B300 | Cisco UCS C880A M8 (8x NVIDIA B300-SXM-270GB, TensorRT) | NVIDIA B300 | 13368 | 99% of FP32 and 99.9% of FP32 (rouge1=42.9865, rouge2=20.1235, rougeL=29.9881). Additionally, for both cases the total generation length of the texts should be more than 90% of the reference (gen_len=8167644) | TTFT/TPOT: 500 ms/30 ms | CNN Dailymail (v3.0.0, max_seq_len=2048) |
**The primary metric on Wan2.2 in Server Scenario is measured in seconds (lower the better).
†Vera Rubin NVL72 results are preview submissions.
MLPerf™ v6.1 Inference Closed Division. NVIDIA platform results from the following entries: 6.1-0008, 6.1-0010, 6.1-0014, 6.1-0015, 6.1-0021, 6.1-0023, 6.1-0025, 6.1-0028, 6.1-0046, 6.1-0064, 6.1-0071, 6.1-0072, 6.1-0073, 6.1-0074, 6.1-0104, 6.1-0106. MLPerf name and logo are trademarks. See
https://mlcommons.org/ for more information.
For MLPerf™ various scenario data, click
here
For MLPerf™ latency constraints, click
here
| Network | Batch Size | Throughput | Efficiency | Latency (ms) | GPU | Server | Container | Precision | Dataset | Framework | GPU Version |
|---|---|---|---|---|---|---|---|---|---|---|---|
| BEVFusion Head | 1 | 2,634 images/sec | 5.52 images/sec/watt | 0.38 | 1x B200 | DGX B200 | 26.08-py3 | INT8 | Synthetic | TensorRT | NVIDIA B200 |
| ControlNet | 4 | 6.28 images/sec | - | 636.55 | 1x B200 | DGX B200 | 26.07-py3 | Mixed | Synthetic | TensorRT | NVIDIA B200 |
| HF Swin Base | 64 | 4,886 samples/sec | 5.34 samples/sec/watt | 13.10 | 1x B200 | DGX B200 | 26.08-py3 | FP8 | Synthetic | TensorRT | NVIDIA B200 |
| HF Swin Large | 128 | 3,127 samples/sec | 3.21 samples/sec/watt | 40.93 | 1x B200 | DGX B200 | 26.08-py3 | FP8 | Synthetic | TensorRT | NVIDIA B200 |
| HF ViT Base | 2048 | 9,366 samples/sec | 9.84 samples/sec/watt | 218.66 | 1x B200 | DGX B200 | 26.07-py3 | FP8 | Synthetic | TensorRT | NVIDIA B200 |
| HF ViT Large | 2048 | 3,364 samples/sec | 3.54 samples/sec/watt | 608.87 | 1x B200 | DGX B200 | 26.08-py3 | FP8 | Synthetic | TensorRT | NVIDIA B200 |
| Stable Diffusion XL | 4 | 2.07 images/sec | - | 1936.19 | 1x B200 | DGX B200 | 26.07-py3 | Mixed | Synthetic | TensorRT | NVIDIA B200 |
| Stable Video Diffusion | 1 | 7.00 videos/min | - | 8570.96 | 1x B200 | DGX B200 | 26.07-py3 | Mixed | Synthetic | TensorRT | NVIDIA B200 |
| Yolo v10 M | 1 | 865 images/sec | 1.07 images/sec/watt | 1.16 | 1x B200 | DGX B200 | 26.08-py3 | INT8 | Synthetic | TensorRT | NVIDIA B200 |
| Yolo v11 M | 1 | 1,079 images/sec | 1.32 images/sec/watt | 0.93 | 1x B200 | DGX B200 | 26.08-py3 | INT8 | Synthetic | TensorRT | NVIDIA B200 |
| Yolo v11 S | 1 | 1,695 images/sec | 2.34 images/sec/watt | 0.59 | 1x B200 | DGX B200 | 26.08-py3 | INT8 | Synthetic | TensorRT | NVIDIA B200 |
HF Swin Base, HF Swin Large, HF ViT Base, HF ViT Large Sequence Length = 384
| Network | Batch Size | Throughput | Efficiency | Latency (ms) | GPU | Server | Container | Precision | Dataset | Framework | GPU Version |
|---|---|---|---|---|---|---|---|---|---|---|---|
| BEVFusion Head | 1 | 1,759 images/sec | 4.57 images/sec/watt | 0.57 | 1x RTX PRO 6000 | Supermicro SYS-521GE-TNRT | 26.07-py3 | FP8 | Synthetic | TensorRT | RTX PRO 6000 BSE |
| ControlNet | 4 | 2.77 images/sec | - | 1442.47 | 1x RTX PRO 6000 | Supermicro SYS-521GE-TNRT | 26.07-py3 | Mixed | Synthetic | TensorRT | RTX PRO 6000 BSE |
| Flux Image Generator | 1 | 0.20 images/sec | - | 5012.06 | 1x RTX PRO 6000 | Supermicro SYS-521GE-TNRT | 26.07-py3 | FP4 | Synthetic | TensorRT | RTX PRO 6000 BSE |
| HF Swin Base | 32 | 2,713 samples/sec | 4.58 samples/sec/watt | 11.79 | 1x RTX PRO 6000 | Supermicro SYS-521GE-TNRT | 26.07-py3 | FP8 | Synthetic | TensorRT | RTX PRO 6000 BSE |
| HF Swin Large | 32 | 1,512 samples/sec | 2.52 samples/sec/watt | 21.17 | 1x RTX PRO 6000 | Supermicro SYS-521GE-TNRT | 26.07-py3 | FP8 | Synthetic | TensorRT | RTX PRO 6000 BSE |
| HF ViT Base | 512 | 3,426 samples/sec | 5.70 samples/sec/watt | 149.45 | 1x RTX PRO 6000 | Supermicro SYS-521GE-TNRT | 26.07-py3 | FP8 | Synthetic | TensorRT | RTX PRO 6000 BSE |
| HF ViT Large | 1024 | 1,165 samples/sec | 2.00 samples/sec/watt | 878.99 | 1x RTX PRO 6000 | Supermicro SYS-521GE-TNRT | 26.07-py3 | FP8 | Synthetic | TensorRT | RTX PRO 6000 BSE |
| Stable Diffusion XL | 4 | 0.72 images/sec | - | 5577.68 | 1x RTX PRO 6000 | Supermicro SYS-521GE-TNRT | 26.07-py3 | Mixed | Synthetic | TensorRT | RTX PRO 6000 BSE |
| Stable Video Diffusion | 1 | 2.82 videos/min | - | 21294.19 | 1x RTX PRO 6000 | Supermicro SYS-521GE-TNRT | 26.07-py3 | Mixed | Synthetic | TensorRT | RTX PRO 6000 BSE |
| Yolo v10 M | 1 | 448 images/sec | 0.83 images/sec/watt | 2.23 | 1x RTX PRO 6000 | Supermicro SYS-521GE-TNRT | 26.07-py3 | FP8 | Synthetic | TensorRT | RTX PRO 6000 BSE |
| Yolo v11 M | 1 | 466 images/sec | 0.86 images/sec/watt | 2.15 | 1x RTX PRO 6000 | Supermicro SYS-521GE-TNRT | 26.07-py3 | FP8 | Synthetic | TensorRT | RTX PRO 6000 BSE |
| Yolo v11 S | 1 | 942 images/sec | 1.79 images/sec/watt | 1.06 | 1x RTX PRO 6000 | Supermicro SYS-521GE-TNRT | 26.07-py3 | FP8 | Synthetic | TensorRT | RTX PRO 6000 BSE |
HF Swin Base, HF Swin Large, HF ViT Base, HF ViT Large Sequence Length = 384
| Network | Batch Size | Throughput | Efficiency | Latency (ms) | GPU | Server | Container | Precision | Dataset | Framework | GPU Version |
|---|---|---|---|---|---|---|---|---|---|---|---|
| BEVFusion Head | 1 | 2,015 images/sec | 6.21 images/sec/watt | 0.50 | 1x H100 | DGX H100 | 26.08-py3 | INT8 | Synthetic | TensorRT | H100 80GB HBM3 |
| ControlNet | 4 | 3.54 images/sec | - | 1129.35 | 1x H100 | DGX H100 | 26.07-py3 | Mixed | Synthetic | TensorRT | H100 80GB HBM3 |
| Flux Image Generator | 1 | 0.20 images/sec | - | 5048.89 | 1x H100 | DGX H100 | 26.07-py3 | FP8 | Synthetic | TensorRT | H100 80GB HBM3 |
| HF Swin Base | 128 | 2,971 samples/sec | 4.40 samples/sec/watt | 43.19 | 1x H100 | DGX H100 | 26.07-py3 | FP8 | Synthetic | TensorRT | H100 80GB HBM3 |
| HF Swin Large | 128 | 1,834 samples/sec | 2.66 samples/sec/watt | 70.02 | 1x H100 | DGX H100 | 26.07-py3 | FP8 | Synthetic | TensorRT | H100 80GB HBM3 |
| HF ViT Base | 2048 | 4,992 samples/sec | 7.20 samples/sec/watt | 410.28 | 1x H100 | DGX H100 | 26.07-py3 | FP8 | Synthetic | TensorRT | H100 80GB HBM3 |
| HF ViT Large | 512 | 1,739 samples/sec | 2.51 samples/sec/watt | 299.82 | 1x H100 | DGX H100 | 26.07-py3 | FP8 | Synthetic | TensorRT | H100 80GB HBM3 |
| Stable Diffusion XL | 4 | 0.98 images/sec | - | 4060.83 | 1x H100 | DGX H100 | 26.07-py3 | Mixed | Synthetic | TensorRT | H100 80GB HBM3 |
| Stable Video Diffusion | 1 | 3.76 videos/min | - | 15964.16 | 1x H100 | DGX H100 | 26.07-py3 | Mixed | Synthetic | TensorRT | H100 80GB HBM3 |
| Yolo v10 M | 1 | 407 images/sec | 0.68 images/sec/watt | 2.47 | 1x H100 | DGX H100 | 26.08-py3 | FP8 | Synthetic | TensorRT | H100 80GB HBM3 |
| Yolo v11 M | 1 | 478 images/sec | 0.79 images/sec/watt | 2.10 | 1x H100 | DGX H100 | 26.07-py3 | FP8 | Synthetic | TensorRT | H100 80GB HBM3 |
| Yolo v11 S | 1 | 948 images/sec | 1.91 images/sec/watt | 1.26 | 1x H100 | DGX H100 | 26.08-py3 | INT8 | Synthetic | TensorRT | H100 80GB HBM3 |
HF Swin Base, HF Swin Large, HF ViT Base, HF ViT Large Sequence Length = 384
| Network | Batch Size | Throughput | Efficiency | Latency (ms) | GPU | Server | Container | Precision | Dataset | Framework | GPU Version |
|---|---|---|---|---|---|---|---|---|---|---|---|
| BEVFusion Head | 1 | 1,953 images/sec | 6.89 images/sec/watt | 0.51 | 1x L40S | Supermicro SYS-521GE-TNRT | 26.08-py3 | INT8 | Synthetic | TensorRT | NVIDIA L40S |
| ControlNet | 4 | 1.59 images/sec | - | 2520.76 | 1x L40S | Supermicro SYS-521GE-TNRT | 26.07-py3 | Mixed | Synthetic | TensorRT | NVIDIA L40S |
| Flux Image Generator | 1 | 0.08 images/sec | - | 12505.61 | 1x L40S | Supermicro SYS-521GE-TNRT | 26.07-py3 | FP8 | Synthetic | TensorRT | NVIDIA L40S |
| HF Swin Base | 32 | 1,392 samples/sec | 4.17 samples/sec/watt | 23.29 | 1x L40S | Supermicro SYS-521GE-TNRT | 26.07-py3 | FP8 | Synthetic | TensorRT | NVIDIA L40S |
| HF Swin Large | 32 | 703 samples/sec | 2.16 samples/sec/watt | 45.80 | 1x L40S | Supermicro SYS-521GE-TNRT | 26.07-py3 | FP8 | Synthetic | TensorRT | NVIDIA L40S |
| HF ViT Base | 1024 | 1,661 samples/sec | 4.81 samples/sec/watt | 616.43 | 1x L40S | Supermicro SYS-521GE-TNRT | 26.07-py3 | FP8 | Synthetic | TensorRT | NVIDIA L40S |
| HF ViT Large | 512 | 592 samples/sec | 1.71 samples/sec/watt | 870.93 | 1x L40S | Supermicro SYS-521GE-TNRT | 26.07-py3 | FP8 | Synthetic | TensorRT | NVIDIA L40S |
| Stable Diffusion XL | 1 | 0.36 images/sec | - | 2738.24 | 1x L40S | Supermicro SYS-521GE-TNRT | 26.07-py3 | Mixed | Synthetic | TensorRT | NVIDIA L40S |
| Stable Video Diffusion | 1 | 1.34 videos/min | - | 44632.28 | 1x L40S | Supermicro SYS-521GE-TNRT | 26.07-py3 | Mixed | Synthetic | TensorRT | NVIDIA L40S |
| Yolo v10 M | 1 | 286 images/sec | 0.83 images/sec/watt | 3.64 | 1x L40S | Supermicro SYS-521GE-TNRT | 26.08-py3 | INT8 | Synthetic | TensorRT | NVIDIA L40S |
| Yolo v11 M | 1 | 328 images/sec | 0.95 images/sec/watt | 3.22 | 1x L40S | Supermicro SYS-521GE-TNRT | 26.08-py3 | INT8 | Synthetic | TensorRT | NVIDIA L40S |
| Yolo v11 S | 1 | 803 images/sec | 2.34 images/sec/watt | 1.35 | 1x L40S | Supermicro SYS-521GE-TNRT | 26.08-py3 | INT8 | Synthetic | TensorRT | NVIDIA L40S |
HF Swin Base, HF Swin Large, HF ViT Base, HF ViT Large Sequence Length = 384
Deploying AI in real-world applications requires training networks to convergence at a specified accuracy. This is the best methodology to test whether AI systems are ready to be deployed in the field to deliver meaningful results.
NVIDIA Riva is an application framework for multimodal conversational AI services that deliver real-performance on GPUs.
Only looking at compute pricing or FLOPs per dollar gives an incomplete view of inference TCO. The most important metric for AI inference TCO is cost per token, or the price-performance actually delivered. GB300 NVL72 delivers AI inference at $0.123 per million tokens at 116 TPS/user interactivity using Dynamo and TensorRT-LLM — the lowest cost per token among major platforms according to SemiAnalysis InferenceX benchmarks as of April 2026.
Metric |
NVIDIA Hopper (HGX H200) |
NVIDIA Blackwell (GB300 NVL72) |
NVIDIA Blackwell Relative to Hopper |
|---|---|---|---|
Cost per GPU per Hour ($) |
$1.41 |
$2.65 |
2x |
FLOP per Dollar (PFLOPS) |
2.8 |
5.6 |
2x |
Tokens per Second per GPU |
90 |
6,000 |
65x |
Tokens per Second per MW |
54K |
2.8M |
50x |
Cost per Million Tokens ($) |
$4.20 |
$0.12 |
35x lower |
GB300 NVL72 delivers AI inference at $0.123 per million tokens at 116 TPS/user interactivity using Dynamo and TensorRT-LLM — the lowest cost per token among major platforms according to SemiAnalysis InferenceX benchmarks as of April 2026.
NVIDIA inference cost per million tokens has improved dramatically across generations: NVIDIA Blackwell Ultra (GB300 NVL72) delivers up to 50x higher throughput per megawatt and up to 35x lower cost per token than NVIDIA Hopper for low-latency agentic workloads, through hardware–software codesign, according to SemiAnalysis InferenceX benchmarks (Q1 2026). Software optimization drives continuous improvement—GB200 token output improved 4x in three months, resulting in a proportional decrease in token cost.
NVIDIA's TensorRT-LLM and Dynamo software stack delivers continuous inference cost improvements without hardware changes. NVIDIA Blackwell B200 cost per million tokens dropped from $0.11 at launch to $0.02 on GPT-OSS-120B within two months, according to SemiAnalysis InferenceX benchmarks as of April 2026—a 5x improvement from software alone. Each TensorRT-LLM release typically delivers throughput gains through kernel fusion, quantization improvements, and scheduling optimizations.