Simulation / Modeling / Design

NVIDIA Releases TensorRT 4

AI-Generated Summary

  • TensorRT 4 accelerates inference workloads such as neural machine translation, recommender systems and speech applications.
  • The release adds new layers for Multilayer Perceptrons and Recurrent Neural Networks to deliver up to 45x higher throughput versus CPU.
  • Models from frameworks including Caffe 2, Chainer, MxNet, Microsoft Cognitive Toolkit and PyTorch can be imported through the ONNX format, achieving 50x faster inference on V100 versus CPU-only for ONNX models parsed in TensorRT.
  • Support extends to NVIDIA DRIVE Xavier for autonomous vehicles, and FP16 custom layers gain a 3x speedup using APIs for Volta Tensor Cores.

Next Steps

Powered by NVIDIA Nemotron. AI-generated content may summarize information incompletely. Verify important information. Learn more

Today we are releasing TensorRT 4 with capabilities for accelerating popular inference applications such as neural machine translation, recommender systems and speech. You also get an easy way to import models from popular deep learning frameworks such as Caffe 2, Chainer, MxNet, Microsoft Cognitive Toolkit and PyTorch through the ONNX format.
TensorRT delivers:

  • Up to 45x higher throughput vs. CPU with new layers for Multilayer Perceptrons (MLP) and Recurrent Neural Networks (RNN)
  • 50x faster inference performance on V100 vs. CPU-only for ONNX models imported with ONNX parser in TensorRT
  • Support for NVIDIA DRIVE Xavier – AI Computer for Autonomous Vehicles
  • 3x Inference speedup for FP16 custom layers with APIs for running on Volta Tensor Cores


Download TensorRT 4 today and try out these exciting new features!
Read more>

Discuss (0)

Tags