Phuong Nguyen

Phuong Nguyen is a senior DL performance engineer at NVIDIA, working on developing the Transformer Engine. Her work spans mixed-precision training and compute-communication overlap for both dense and MoE models across JAX and PyTorch. She is also a primary contributor to NCCL EP–a communication backend purpose-built for MoE token routing–and integrates it into Transformer Engine. Before NVIDIA, she worked in high-performance computing, optimizing kernels for sparse linear algebra and scientific computing.
Avatar photo

Posts by Phuong Nguyen

MLOps

Accelerating Dropless MoE Training in JAX with NVIDIA Transformer Engine

Mixture of experts (MoE) has become one of the defining architectural trends in large-scale AI model training. DeepSeek, Qwen, and Mixtral are examples of MoE... 12 MIN READ