Networking / Communications

SC20 Demo: Maximizing Performance for Distributed Machine Learning and Deep Learning with SHARP

AI-Generated Summary

  • NVIDIA Mellanox SHARP offloads collective operations from the CPU to the network using in-network computing in the NVIDIA Mellanox Quantum switch.
  • The protocol eliminates the need to send data multiple times between endpoints, decreasing the amount of data traversing the network as aggregation nodes are reached.
  • This approach dramatically reduces collective operations time for distributed machine learning workloads.

Next Steps

  • Learn more about SHARP for AI workloads.
  • Explore Magnum IO for accelerated data center I/O.
  • View all SC20 Demos from NVIDIA.
Powered by NVIDIA Nemotron. AI-generated content may summarize information incompletely. Verify important information. Learn more

Today’s modern-day machine learning data centers require complex computations and fast, efficient data delivery. The NVIDIA Mellanox Scalable Hierarchical Aggregation and Reduction Protocol (SHARP) takes advantage of the in-network computing capabilities in the NVIDIA Mellanox Quantum switch, dramatically improving the performance of distributed machine learning workloads.

SHARP technology improves upon the performance of MPI and Machine Learning collective operations by offloading collective operations from the CPU to the network and eliminating the need to send data multiple times between endpoints.

This innovative approach decreases the amount of data traversing the network as aggregation nodes are reached, and dramatically reduces collective operations time. 

Magnum IO> 

Learn More about SHARP

View all SC20 Demos> 

Discuss (0)

Tags