AI Inference Performance Benchmarks

This is NVIDIA's Data Center Deep Learning Product Performance Hub — a centralized resource for reproducible AI performance benchmarks across NVIDIA's latest data center GPUs.

The days of raw speed being the only metric that matters are behind us. Now it’s about throughput, efficiency, and economics at scale. As AI evolves from providing one-shot answers to engaging in multi-step reasoning, the demand for inference and its underlying economics is increasing.. This shift significantly boosts compute demand due to the generation of far more tokens per query. Metrics such as tokens per watt, cost per million tokens, and tokens per second per user are crucial alongside throughput.

For power-limited AI factories, NVIDIA's continuous software improvements translate into higher token revenue over time, underscoring the importance of our technological advancements.

Pareto curves illustrate how NVIDIA Vera Rubin and  provides the best balance across the full spectrum of production priorities, including:

  • Cost

  • Energy efficiency

  • Throughput

  • Responsiveness

Optimizing systems for a single scenario can limit deployment flexibility,‌ leading to inefficiencies at other points on the curve. NVIDIA’s full-stack design approach ensures efficiency and value across multiple real-life production scenarios. Blackwell’s leadership stems from its extreme hardware-software co-design, embodying a full-stack architecture built for speed, efficiency, and scalability.

Explore the methodology used to obtain these results and learn how to replicate the tests by executing Benchmarking Recipes yourself.

MLPerf Inference v6.1 Performance Benchmarks

Offline Scenario, Closed Division

Network Throughput GPU Server GPU Version QSL Size Target Accuracy Dataset
DeepSeek R1 1,183,327 tokens/sec 72x VR200† NVIDIA Vera Rubin NVL72 (72x VR200-288GB_aarch64) NVIDIA Vera Rubin NVL72 4388 99% of FP16 (exact match 81.9132%) mlperf_deepseek_r1
2,705,130 tokens/sec 288x GB300 NVIDIA GB300 NVL72 (72x GB300-288GB_aarch64, TensorRT) NVIDIA GB300 NVL72 4388 99% of FP16 (exact match 81.9132%) mlperf_deepseek_r1
2,091,190 tokens/sec 288x GB200 Azure GB200 NVL72 (288x GB200-186GB_aarch64, TensorRT) NVIDIA GB200 NVL72 4388 99% of FP16 (exact match 81.9132%) mlperf_deepseek_r1
137,067 tokens/sec 16x B300 Cisco_UCS_B300x16_G200_4x2_RO NVIDIA B300 4388 99% of FP16 (exact match 81.9132%) mlperf_deepseek_r1
58,913 tokens/sec 8x B200 CoreWeave B200 SXM (8x B200-SXM-180GB) NVIDIA B200 4388 99% of FP16 (exact match 81.9132%) mlperf_deepseek_r1
gpt-oss 120B 1,197,710 tokens/sec 72x GB300 CoreWeave GB300 NVL72 (72x GB300-288GB_aarch64) NVIDIA GB300 NVL72 6396 99% of 83.13% AIME25, GPQA Diamond, LiveCodeBench v6
911,566 tokens/sec 72x GB200 CoreWeave GB200 NVL72 (72x GB200-186GB_aarch64) NVIDIA GB200 NVL72 6396 99% of 83.13% AIME25, GPQA Diamond, LiveCodeBench v6
220,527 tokens/sec 16x B300 Cisco_UCS_B300x16_G200_4x2_RO NVIDIA B300 6396 99% of 83.13% AIME25, GPQA Diamond, LiveCodeBench v6
91,487 tokens/sec 8x B200 CoreWeave B200 SXM (8x B200-SXM-180GB) NVIDIA B200 6396 99% of 83.13% AIME25, GPQA Diamond, LiveCodeBench v6
Qwen3-VL 235B 2,393 samples/sec 72x VR200† NVIDIA Vera Rubin NVL72 (72x VR200-288GB_aarch64) NVIDIA Vera Rubin NVL72 48289 99% of BF16 (Category Hierarchical F1 Score >= 0.7824) Shopify Product Catalogue
1,305 samples/sec 72x GB300 NVIDIA GB300 NVL72 (72x GB300-288GB_aarch64, TensorRT) NVIDIA GB300 NVL72 48289 99% of BF16 (Category Hierarchical F1 Score >= 0.7824) Shopify Product Catalogue
230 samples/sec 16x GB200 NVIDIA GB200 NVL72 (16x GB200-186GB_aarch64, Dynamo) NVIDIA GB200 NVL72 48289 99% of BF16 (Category Hierarchical F1 Score >= 0.7824) Shopify Product Catalogue
132 samples/sec 8x B300 G894-ZD3-AAX7 NVIDIA B300 48289 99% of BF16 (Category Hierarchical F1 Score >= 0.7824) Shopify Product Catalogue
102 samples/sec 8x B200 Lambda Cloud 8x B200-SXM-180GB NVIDIA B200 48289 99% of BF16 (Category Hierarchical F1 Score >= 0.7824) Shopify Product Catalogue
Llama2 70B 1,136,100 tokens/sec 72x GB300 CoreWeave GB300 NVL72 (72x GB300-288GB_aarch64) NVIDIA GB300 NVL72 24576 99.9% of FP32 (rouge1=44.4312, rouge2=22.0352, rougeL=28.6162). Additionally, the generation length of the tokens per sample should be more than 90% of the reference (tokens_per_sample=294.45) OpenOrca (max_seq_len=1024)
894,002 tokens/sec 72x GB200 CoreWeave GB200 NVL72 (72x GB200-186GB_aarch64) NVIDIA GB200 NVL72 24576 99.9% of FP32 (rouge1=44.4312, rouge2=22.0352, rougeL=28.6162). Additionally, the generation length of the tokens per sample should be more than 90% of the reference (tokens_per_sample=294.45) OpenOrca (max_seq_len=1024)
227,570 tokens/sec 16x B300 Cisco_UCS_B300x16_G200_4x2_RO NVIDIA B300 24576 99.9% of FP32 (rouge1=44.4312, rouge2=22.0352, rougeL=28.6162). Additionally, the generation length of the tokens per sample should be more than 90% of the reference (tokens_per_sample=294.45) OpenOrca (max_seq_len=1024)
102,703 tokens/sec 8x B200 CoreWeave B200 SXM (8x B200-SXM-180GB) NVIDIA B200 24576 99.9% of FP32 (rouge1=44.4312, rouge2=22.0352, rougeL=28.6162). Additionally, the generation length of the tokens per sample should be more than 90% of the reference (tokens_per_sample=294.45) OpenOrca (max_seq_len=1024)
Llama3.1 8B 171,114 tokens/sec 8x B300 Cisco UCS C880A M8 (8x NVIDIA B300-SXM-270GB, TensorRT) NVIDIA B300 13368 99% of FP32 and 99.9% of FP32 (rouge1=42.9865, rouge2=20.1235, rougeL=29.9881). Additionally, for both cases the total generation length of the texts should be more than 90% of the reference (gen_len=8167644) CNN Dailymail (v3.0.0, max_seq_len=2048)
Whisper 53,267 samples/sec 8x B300 NVIDIA DGX B300 (8x B300-SXM-270GB, TensorRT) NVIDIA B300 1633 99% of FP32 and 99.9% of FP32 (WER=2.0671%) LibriSpeech
Wan2.2 0.655 samples/sec 72x GB300 NVIDIA GB300 NVL72 (72x GB300-288GB_aarch64, TensorRT) NVIDIA GB300 NVL72 248 99% of BF16 (VBench score >= 69.7752) VBench prompts
0.496 samples/sec 72x GB200 NVIDIA GB200 NVL72 (72x GB200-186GB_aarch64, TensorRT) NVIDIA GB200 NVL72 248 99% of BF16 (VBench score >= 69.7752) VBench prompts
0.078 samples/sec 8x B300 NVIDIA DGX B300 (8x B300-SXM-270GB, TensorRT) NVIDIA B300 248 99% of BF16 (VBench score >= 69.7752) VBench prompts

Server Scenario - Closed Division

Network Throughput GPU Server GPU Version QSL Size Target Accuracy MLPerf Server Latency
Constraints (ms)
Dataset
DeepSeek R1 1,175,890 tokens/sec 72x VR200† NVIDIA Vera Rubin NVL72 (72x VR200-288GB_aarch64) NVIDIA Vera Rubin NVL72 4388 99% of FP16 (exact match 81.9132%) TTFT/TPOT: 2000 ms/80 ms mlperf_deepseek_r1
2,028,030 tokens/sec 288x GB300 NVIDIA GB300 NVL72 (72x GB300-288GB_aarch64, TensorRT) NVIDIA GB300 NVL72 4388 99% of FP16 (exact match 81.9132%) TTFT/TPOT: 2000 ms/80 ms mlperf_deepseek_r1
1,598,510 tokens/sec 288x GB200 Azure GB200 NVL72 (288x GB200-186GB_aarch64, TensorRT) NVIDIA GB200 NVL72 4388 99% of FP16 (exact match 81.9132%) TTFT/TPOT: 2000 ms/80 ms mlperf_deepseek_r1
88,916 tokens/sec 16x B300 Cisco_UCS_B300x16_G200_4x2_RO NVIDIA B300 4388 99% of FP16 (exact match 81.9132%) TTFT/TPOT: 2000 ms/80 ms mlperf_deepseek_r1
56,121 tokens/sec 8x B200 CoreWeave B200 SXM (8x B200-SXM-180GB) NVIDIA B200 4388 99% of FP16 (exact match 81.9132%) TTFT/TPOT: 2000 ms/80 ms mlperf_deepseek_r1
gpt-oss 120B 1,160,490 tokens/sec 72x GB300 CoreWeave GB300 NVL72 (72x GB300-288GB_aarch64) NVIDIA GB300 NVL72 6396 99% of 83.13% TTFT/TPOT: 3000 ms/80 ms AIME25, GPQA Diamond, LiveCodeBench v6
901,058 tokens/sec 72x GB200 CoreWeave GB200 NVL72 (72x GB200-186GB_aarch64) NVIDIA GB200 NVL72 6396 99% of 83.13% TTFT/TPOT: 3000 ms/80 ms AIME25, GPQA Diamond, LiveCodeBench v6
214,774 tokens/sec 16x B300 Cisco_UCS_B300x16_G200_4x2_RO NVIDIA B300 6396 99% of 83.13% TTFT/TPOT: 3000 ms/80 ms AIME25, GPQA Diamond, LiveCodeBench v6
90,205 tokens/sec 8x B200 CoreWeave B200 SXM (8x B200-SXM-180GB) NVIDIA B200 6396 99% of 83.13% TTFT/TPOT: 3000 ms/80 ms AIME25, GPQA Diamond, LiveCodeBench v6
Qwen3-VL 235B 2,323 queries/sec 72x VR200† NVIDIA Vera Rubin NVL72 (72x VR200-288GB_aarch64) NVIDIA Vera Rubin NVL72 48289 99% of BF16 (Category Hierarchical F1 Score >= 0.7824) 12 s Shopify Product Catalogue
1,210 queries/sec 72x GB300 NVIDIA GB300 NVL72 (72x GB300-288GB_aarch64, TensorRT) NVIDIA GB300 NVL72 48289 99% of BF16 (Category Hierarchical F1 Score >= 0.7824) 12 s Shopify Product Catalogue
199 queries/sec 16x GB200 NVIDIA GB200 NVL72 (16x GB200-186GB_aarch64, Dynamo) NVIDIA GB200 NVL72 48289 99% of BF16 (Category Hierarchical F1 Score >= 0.7824) 12 s Shopify Product Catalogue
116 queries/sec 8x B300 G894-ZD3-AAX7 NVIDIA B300 48289 99% of BF16 (Category Hierarchical F1 Score >= 0.7824) 12 s Shopify Product Catalogue
69.415 queries/sec 8x B200 Lambda Cloud 8x B200-SXM-180GB NVIDIA B200 48289 99% of BF16 (Category Hierarchical F1 Score >= 0.7824) 12 s Shopify Product Catalogue
Llama2 70B 944,902 tokens/sec 72x GB300 CoreWeave GB300 NVL72 (72x GB300-288GB_aarch64) NVIDIA GB300 NVL72 24576 99.9% of FP32 (rouge1=44.4312, rouge2=22.0352, rougeL=28.6162). Additionally, the generation length of the tokens per sample should be more than 90% of the reference (tokens_per_sample=294.45) TTFT/TPOT: 2000 ms/200 ms OpenOrca (max_seq_len=1024)
910,621 tokens/sec 72x GB200 CoreWeave GB200 NVL72 (72x GB200-186GB_aarch64) NVIDIA GB200 NVL72 24576 99.9% of FP32 (rouge1=44.4312, rouge2=22.0352, rougeL=28.6162). Additionally, the generation length of the tokens per sample should be more than 90% of the reference (tokens_per_sample=294.45) TTFT/TPOT: 2000 ms/200 ms OpenOrca (max_seq_len=1024)
179,597 tokens/sec 16x B300 Cisco_UCS_B300x16_G200_4x2_RO NVIDIA B300 24576 99.9% of FP32 (rouge1=44.4312, rouge2=22.0352, rougeL=28.6162). Additionally, the generation length of the tokens per sample should be more than 90% of the reference (tokens_per_sample=294.45) TTFT/TPOT: 2000 ms/200 ms OpenOrca (max_seq_len=1024)
102,398 tokens/sec 8x B200 CoreWeave B200 SXM (8x B200-SXM-180GB) NVIDIA B200 24576 99.9% of FP32 (rouge1=44.4312, rouge2=22.0352, rougeL=28.6162). Additionally, the generation length of the tokens per sample should be more than 90% of the reference (tokens_per_sample=294.45) TTFT/TPOT: 2000 ms/200 ms OpenOrca (max_seq_len=1024)
Llama3.1 8B 173,950 tokens/sec 8x B300 Cisco UCS C880A M8 (8x NVIDIA B300-SXM-270GB, TensorRT) NVIDIA B300 13368 99% of FP32 and 99.9% of FP32 (rouge1=42.9865, rouge2=20.1235, rougeL=29.9881). Additionally, for both cases the total generation length of the texts should be more than 90% of the reference (gen_len=8167644) TTFT/TPOT: 2000 ms/100 ms CNN Dailymail (v3.0.0, max_seq_len=2048)
Wan2.2** 5.73 seconds 72x GB200 NVIDIA GB200 NVL72 (72x GB200-186GB_aarch64, TensorRT) NVIDIA GB200 NVL72 248 99% of BF16 (VBench score >= 69.7752) N/A VBench prompts
5.81 seconds 72x GB300 NVIDIA GB300 NVL72 (72x GB300-288GB_aarch64, TensorRT) NVIDIA GB300 NVL72 248 99% of BF16 (VBench score >= 69.7752) N/A VBench prompts
17.98 seconds 8x B300 NVIDIA DGX B300 (8x B300-SXM-270GB, TensorRT) NVIDIA B300 248 99% of BF16 (VBench score >= 69.7752) N/A VBench prompts

Interactive Scenario - Closed Division

Network Throughput GPU Server GPU Version QSL Size Target Accuracy MLPerf Server Latency
Constraints (ms)
Dataset
DeepSeek R1 652,750 tokens/sec 72x VR200† NVIDIA Vera Rubin NVL72 (72x VR200-288GB_aarch64) NVIDIA Vera Rubin NVL72 4388 99% of FP16 (exact match 81.9132%) TTFT/TPOT: 1500 ms/15 ms mlperf_deepseek_r1
260,098 tokens/sec 72x GB300 Azure GB300 NVL72 (72x GB300-288GB_aarch64 TensorRT) NVIDIA GB300 NVL72 4388 99% of FP16 (exact match 81.9132%) TTFT/TPOT: 1500 ms/15 ms mlperf_deepseek_r1
240,274 tokens/sec 72x GB200 NVIDIA GB200 NVL72 (72x GB200-186GB_aarch64, TensorRT) NVIDIA GB200 NVL72 4388 99% of FP16 (exact match 81.9132%) TTFT/TPOT: 1500 ms/15 ms mlperf_deepseek_r1
gpt-oss 120B 669,305 tokens/sec 72x GB300 NVIDIA GB300 NVL72 (72x GB300-288GB_aarch64, TensorRT) NVIDIA GB300 NVL72 6396 99% of 83.13% TTFT/TPOT: 2000 ms/20 ms AIME25, GPQA Diamond, LiveCodeBench v6
625,205 tokens/sec 72x GB200 NVIDIA GB200 NVL72 (72x GB200-186GB_aarch64, TensorRT) NVIDIA GB200 NVL72 6396 99% of 83.13% TTFT/TPOT: 2000 ms/20 ms AIME25, GPQA Diamond, LiveCodeBench v6
Qwen3-VL 235B 1,307 queries/sec 72x VR200† NVIDIA Vera Rubin NVL72 (72x VR200-288GB_aarch64) NVIDIA Vera Rubin NVL72 48289 99% of BF16 (Category Hierarchical F1 Score >= 0.7824) 1.5 s Shopify Product Catalogue
349 queries/sec 72x GB300 NVIDIA GB300 NVL72 (72x GB300-288GB_aarch64, TensorRT) NVIDIA GB300 NVL72 48289 99% of BF16 (Category Hierarchical F1 Score >= 0.7824) 1.5 s Shopify Product Catalogue
39.851 queries/sec 16x GB200 NVIDIA GB200 NVL72 (16x GB200-186GB_aarch64, Dynamo) NVIDIA GB200 NVL72 48289 99% of BF16 (Category Hierarchical F1 Score >= 0.7824) 1.5 s Shopify Product Catalogue
12.882 queries/sec 8x B300 G894-ZD3-AAX7 NVIDIA B300 48289 99% of BF16 (Category Hierarchical F1 Score >= 0.7824) 1.5 s Shopify Product Catalogue
Llama3.1 8B 129,168 tokens/sec 8x B300 Cisco UCS C880A M8 (8x NVIDIA B300-SXM-270GB, TensorRT) NVIDIA B300 13368 99% of FP32 and 99.9% of FP32 (rouge1=42.9865, rouge2=20.1235, rougeL=29.9881). Additionally, for both cases the total generation length of the texts should be more than 90% of the reference (gen_len=8167644) TTFT/TPOT: 500 ms/30 ms CNN Dailymail (v3.0.0, max_seq_len=2048)

**The primary metric on Wan2.2 in Server Scenario is measured in seconds (lower the better).
†Vera Rubin NVL72 results are preview submissions.
MLPerf™ v6.1 Inference Closed Division. NVIDIA platform results from the following entries: 6.1-0008, 6.1-0010, 6.1-0014, 6.1-0015, 6.1-0021, 6.1-0023, 6.1-0025, 6.1-0028, 6.1-0046, 6.1-0064, 6.1-0071, 6.1-0072, 6.1-0073, 6.1-0074, 6.1-0104, 6.1-0106. MLPerf name and logo are trademarks. See https://mlcommons.org/ for more information.
For MLPerf™ various scenario data, click here
For MLPerf™ latency constraints, click here

Inference Performance of NVIDIA Data Center Products

B200 Inference Performance

Network Batch Size Throughput Efficiency Latency (ms) GPU Server Container Precision Dataset Framework GPU Version
BEVFusion Head 1 2,634 images/sec 5.52 images/sec/watt 0.38 1x B200 DGX B200 26.08-py3 INT8 Synthetic TensorRT NVIDIA B200
ControlNet 4 6.28 images/sec - 636.55 1x B200 DGX B200 26.07-py3 Mixed Synthetic TensorRT NVIDIA B200
HF Swin Base 64 4,886 samples/sec 5.34 samples/sec/watt 13.10 1x B200 DGX B200 26.08-py3 FP8 Synthetic TensorRT NVIDIA B200
HF Swin Large 128 3,127 samples/sec 3.21 samples/sec/watt 40.93 1x B200 DGX B200 26.08-py3 FP8 Synthetic TensorRT NVIDIA B200
HF ViT Base 2048 9,366 samples/sec 9.84 samples/sec/watt 218.66 1x B200 DGX B200 26.07-py3 FP8 Synthetic TensorRT NVIDIA B200
HF ViT Large 2048 3,364 samples/sec 3.54 samples/sec/watt 608.87 1x B200 DGX B200 26.08-py3 FP8 Synthetic TensorRT NVIDIA B200
Stable Diffusion XL 4 2.07 images/sec - 1936.19 1x B200 DGX B200 26.07-py3 Mixed Synthetic TensorRT NVIDIA B200
Stable Video Diffusion 1 7.00 videos/min - 8570.96 1x B200 DGX B200 26.07-py3 Mixed Synthetic TensorRT NVIDIA B200
Yolo v10 M 1 865 images/sec 1.07 images/sec/watt 1.16 1x B200 DGX B200 26.08-py3 INT8 Synthetic TensorRT NVIDIA B200
Yolo v11 M 1 1,079 images/sec 1.32 images/sec/watt 0.93 1x B200 DGX B200 26.08-py3 INT8 Synthetic TensorRT NVIDIA B200
Yolo v11 S 1 1,695 images/sec 2.34 images/sec/watt 0.59 1x B200 DGX B200 26.08-py3 INT8 Synthetic TensorRT NVIDIA B200

HF Swin Base, HF Swin Large, HF ViT Base, HF ViT Large Sequence Length = 384

RTX PRO 6000 Blackwell Server Edition Inference Performance

Network Batch Size Throughput Efficiency Latency (ms) GPU Server Container Precision Dataset Framework GPU Version
BEVFusion Head 1 1,759 images/sec 4.57 images/sec/watt 0.57 1x RTX PRO 6000 Supermicro SYS-521GE-TNRT 26.07-py3 FP8 Synthetic TensorRT RTX PRO 6000 BSE
ControlNet 4 2.77 images/sec - 1442.47 1x RTX PRO 6000 Supermicro SYS-521GE-TNRT 26.07-py3 Mixed Synthetic TensorRT RTX PRO 6000 BSE
Flux Image Generator 1 0.20 images/sec - 5012.06 1x RTX PRO 6000 Supermicro SYS-521GE-TNRT 26.07-py3 FP4 Synthetic TensorRT RTX PRO 6000 BSE
HF Swin Base 32 2,713 samples/sec 4.58 samples/sec/watt 11.79 1x RTX PRO 6000 Supermicro SYS-521GE-TNRT 26.07-py3 FP8 Synthetic TensorRT RTX PRO 6000 BSE
HF Swin Large 32 1,512 samples/sec 2.52 samples/sec/watt 21.17 1x RTX PRO 6000 Supermicro SYS-521GE-TNRT 26.07-py3 FP8 Synthetic TensorRT RTX PRO 6000 BSE
HF ViT Base 512 3,426 samples/sec 5.70 samples/sec/watt 149.45 1x RTX PRO 6000 Supermicro SYS-521GE-TNRT 26.07-py3 FP8 Synthetic TensorRT RTX PRO 6000 BSE
HF ViT Large 1024 1,165 samples/sec 2.00 samples/sec/watt 878.99 1x RTX PRO 6000 Supermicro SYS-521GE-TNRT 26.07-py3 FP8 Synthetic TensorRT RTX PRO 6000 BSE
Stable Diffusion XL 4 0.72 images/sec - 5577.68 1x RTX PRO 6000 Supermicro SYS-521GE-TNRT 26.07-py3 Mixed Synthetic TensorRT RTX PRO 6000 BSE
Stable Video Diffusion 1 2.82 videos/min - 21294.19 1x RTX PRO 6000 Supermicro SYS-521GE-TNRT 26.07-py3 Mixed Synthetic TensorRT RTX PRO 6000 BSE
Yolo v10 M 1 448 images/sec 0.83 images/sec/watt 2.23 1x RTX PRO 6000 Supermicro SYS-521GE-TNRT 26.07-py3 FP8 Synthetic TensorRT RTX PRO 6000 BSE
Yolo v11 M 1 466 images/sec 0.86 images/sec/watt 2.15 1x RTX PRO 6000 Supermicro SYS-521GE-TNRT 26.07-py3 FP8 Synthetic TensorRT RTX PRO 6000 BSE
Yolo v11 S 1 942 images/sec 1.79 images/sec/watt 1.06 1x RTX PRO 6000 Supermicro SYS-521GE-TNRT 26.07-py3 FP8 Synthetic TensorRT RTX PRO 6000 BSE

HF Swin Base, HF Swin Large, HF ViT Base, HF ViT Large Sequence Length = 384

H100 Inference Performance

Network Batch Size Throughput Efficiency Latency (ms) GPU Server Container Precision Dataset Framework GPU Version
BEVFusion Head 1 2,015 images/sec 6.21 images/sec/watt 0.50 1x H100 DGX H100 26.08-py3 INT8 Synthetic TensorRT H100 80GB HBM3
ControlNet 4 3.54 images/sec - 1129.35 1x H100 DGX H100 26.07-py3 Mixed Synthetic TensorRT H100 80GB HBM3
Flux Image Generator 1 0.20 images/sec - 5048.89 1x H100 DGX H100 26.07-py3 FP8 Synthetic TensorRT H100 80GB HBM3
HF Swin Base 128 2,971 samples/sec 4.40 samples/sec/watt 43.19 1x H100 DGX H100 26.07-py3 FP8 Synthetic TensorRT H100 80GB HBM3
HF Swin Large 128 1,834 samples/sec 2.66 samples/sec/watt 70.02 1x H100 DGX H100 26.07-py3 FP8 Synthetic TensorRT H100 80GB HBM3
HF ViT Base 2048 4,992 samples/sec 7.20 samples/sec/watt 410.28 1x H100 DGX H100 26.07-py3 FP8 Synthetic TensorRT H100 80GB HBM3
HF ViT Large 512 1,739 samples/sec 2.51 samples/sec/watt 299.82 1x H100 DGX H100 26.07-py3 FP8 Synthetic TensorRT H100 80GB HBM3
Stable Diffusion XL 4 0.98 images/sec - 4060.83 1x H100 DGX H100 26.07-py3 Mixed Synthetic TensorRT H100 80GB HBM3
Stable Video Diffusion 1 3.76 videos/min - 15964.16 1x H100 DGX H100 26.07-py3 Mixed Synthetic TensorRT H100 80GB HBM3
Yolo v10 M 1 407 images/sec 0.68 images/sec/watt 2.47 1x H100 DGX H100 26.08-py3 FP8 Synthetic TensorRT H100 80GB HBM3
Yolo v11 M 1 478 images/sec 0.79 images/sec/watt 2.10 1x H100 DGX H100 26.07-py3 FP8 Synthetic TensorRT H100 80GB HBM3
Yolo v11 S 1 948 images/sec 1.91 images/sec/watt 1.26 1x H100 DGX H100 26.08-py3 INT8 Synthetic TensorRT H100 80GB HBM3

HF Swin Base, HF Swin Large, HF ViT Base, HF ViT Large Sequence Length = 384

L40S Inference Performance

Network Batch Size Throughput Efficiency Latency (ms) GPU Server Container Precision Dataset Framework GPU Version
BEVFusion Head 1 1,953 images/sec 6.89 images/sec/watt 0.51 1x L40S Supermicro SYS-521GE-TNRT 26.08-py3 INT8 Synthetic TensorRT NVIDIA L40S
ControlNet 4 1.59 images/sec - 2520.76 1x L40S Supermicro SYS-521GE-TNRT 26.07-py3 Mixed Synthetic TensorRT NVIDIA L40S
Flux Image Generator 1 0.08 images/sec - 12505.61 1x L40S Supermicro SYS-521GE-TNRT 26.07-py3 FP8 Synthetic TensorRT NVIDIA L40S
HF Swin Base 32 1,392 samples/sec 4.17 samples/sec/watt 23.29 1x L40S Supermicro SYS-521GE-TNRT 26.07-py3 FP8 Synthetic TensorRT NVIDIA L40S
HF Swin Large 32 703 samples/sec 2.16 samples/sec/watt 45.80 1x L40S Supermicro SYS-521GE-TNRT 26.07-py3 FP8 Synthetic TensorRT NVIDIA L40S
HF ViT Base 1024 1,661 samples/sec 4.81 samples/sec/watt 616.43 1x L40S Supermicro SYS-521GE-TNRT 26.07-py3 FP8 Synthetic TensorRT NVIDIA L40S
HF ViT Large 512 592 samples/sec 1.71 samples/sec/watt 870.93 1x L40S Supermicro SYS-521GE-TNRT 26.07-py3 FP8 Synthetic TensorRT NVIDIA L40S
Stable Diffusion XL 1 0.36 images/sec - 2738.24 1x L40S Supermicro SYS-521GE-TNRT 26.07-py3 Mixed Synthetic TensorRT NVIDIA L40S
Stable Video Diffusion 1 1.34 videos/min - 44632.28 1x L40S Supermicro SYS-521GE-TNRT 26.07-py3 Mixed Synthetic TensorRT NVIDIA L40S
Yolo v10 M 1 286 images/sec 0.83 images/sec/watt 3.64 1x L40S Supermicro SYS-521GE-TNRT 26.08-py3 INT8 Synthetic TensorRT NVIDIA L40S
Yolo v11 M 1 328 images/sec 0.95 images/sec/watt 3.22 1x L40S Supermicro SYS-521GE-TNRT 26.08-py3 INT8 Synthetic TensorRT NVIDIA L40S
Yolo v11 S 1 803 images/sec 2.34 images/sec/watt 1.35 1x L40S Supermicro SYS-521GE-TNRT 26.08-py3 INT8 Synthetic TensorRT NVIDIA L40S

HF Swin Base, HF Swin Large, HF ViT Base, HF ViT Large Sequence Length = 384

View More Performance Data

Training to Convergence

Deploying AI in real-world applications requires training networks to convergence at a specified accuracy. This is the best methodology to test whether AI systems are ready to be deployed in the field to deliver meaningful results.

AI Pipeline

NVIDIA Riva is an application framework for multimodal conversational AI services that deliver real-performance on GPUs.

Frequently Asked Questions About Performance of NVIDIA Data Center Deep Learning Products

Only looking at compute pricing or FLOPs per dollar gives an incomplete view of inference TCO. The most important metric for AI inference TCO is cost per token, or the price-performance actually delivered. GB300 NVL72 delivers AI inference at $0.123 per million tokens at 116 TPS/user interactivity using Dynamo and TensorRT-LLM — the lowest cost per token among major platforms according to SemiAnalysis InferenceX benchmarks as of April 2026.

Metric
NVIDIA Hopper (HGX H200)
NVIDIA Blackwell (GB300 NVL72)
NVIDIA Blackwell Relative to Hopper
Cost per GPU per Hour ($)
$1.41
$2.65
2x
FLOP per Dollar (PFLOPS)
2.8
5.6
2x
Tokens per Second per GPU
90
6,000
65x
Tokens per Second per MW
54K
2.8M
50x
Cost per Million Tokens ($)
$4.20
$0.12
35x lower


GB300 NVL72 delivers AI inference at $0.123 per million tokens at 116 TPS/user interactivity using Dynamo and TensorRT-LLM — the lowest cost per token among major platforms according to SemiAnalysis InferenceX benchmarks as of April 2026.

NVIDIA inference cost per million tokens has improved dramatically across generations: NVIDIA Blackwell Ultra (GB300 NVL72) delivers up to 50x higher throughput per megawatt and up to 35x lower cost per token than NVIDIA Hopper for low-latency agentic workloads, through hardware–software codesign, according to SemiAnalysis InferenceX benchmarks (Q1 2026). Software optimization drives continuous improvement—GB200 token output improved 4x in three months, resulting in a proportional decrease in token cost.

NVIDIA's TensorRT-LLM and Dynamo software stack delivers continuous inference cost improvements without hardware changes. NVIDIA Blackwell B200 cost per million tokens dropped from $0.11 at launch to $0.02 on GPT-OSS-120B within two months, according to SemiAnalysis InferenceX benchmarks as of April 2026—a 5x improvement from software alone. Each TensorRT-LLM release typically delivers throughput gains through kernel fusion, quantization improvements, and scheduling optimizations.