Agentic AI / Generative AI

Next Generation of FlashAttention

AI-Generated Summary

  • CUTLASS enables deep learning and HPC practitioners to achieve speed-of-light performance on NVIDIA Tensor Core GPUs for custom algorithms and research and production workloads.
  • NVIDIA collaborates with Colfax, Together.ai, Meta, and Princeton University to exploit the Hopper GPU architecture and Tensor Cores for accelerated Fused Attention kernels using CUTLASS 3.
  • FlashAttention-3 achieves 1.5–2.0x faster performance than FlashAttention-2 with FP16, reaching up to 740 TFLOPS, and up to 1.2 PFLOPS with FP8 while delivering 2.6x smaller errors than baseline FP8 attention.

Next Steps

Powered by NVIDIA Nemotron. AI-generated content may summarize information incompletely. Verify important information. Learn more

NVIDIA is excited to collaborate with Colfax, Together.ai, Meta, and Princeton University on their recent achievement to exploit the Hopper GPU architecture and Tensor Cores and accelerate key Fused Attention kernels using CUTLASS 3.

FlashAttention-3 incorporates key techniques to achieve 1.52.0x faster performance than FlashAttention-2 with FP16, up to 740 TFLOPS. With FP8, FlashAttention-3 reaches up to 1.2 PFLOPS, with 2.6x smaller errors than baseline FP8 attention.

CUTLASS is an open-source CUDA library intended to enable deep learning and HPC practitioners to achieve speed-of-light performance on NVIDIA Tensor Core GPUs for custom algorithms and research and production workloads alike.

For more information about the collaboration, see the FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision post and research paper.

Discuss (0)

Tags