Phuong Nguyen

Phuong Nguyen is a senior DL performance engineer at NVIDIA, working on developing the Transformer Engine. Her work spans mixed-precision training and compute-communication overlap for both dense and MoE models across JAX and PyTorch. She is also a primary contributor to NCCL EP–a communication backend purpose-built for MoE token routing–and integrates it into Transformer Engine. Before NVIDIA, she worked in high-performance computing, optimizing kernels for sparse linear algebra and scientific computing.

Posts by Phuong Nguyen