Daisy Chu

Daisy Chu is a senior systems software engineer on the NVIDIA TensorRT team, specializing in multi-device architectures. Her work centers on building production-grade inference systems, with an emphasis on performance optimization, correctness validation, and scalable execution across single- and multi-GPU environments. Daisy is instrumental in enabling efficient multi-GPU inference for large language and multimodal models, ensuring high scalability and robustness. She holds a master’s degree in Computer Science from the University of Illinois Urbana-Champaign.
Avatar photo

Posts by Daisy Chu

Agentic AI / Generative AI

Simplifying Model Serving Across Multiple GPUs with NVIDIA TensorRT Multi-Device Integration in NVIDIA Dynamo-Triton

The compute and memory demands of generative AI increasingly exceed what a single GPU can provide. NVIDIA TensorRT multi-device inference is a new capability... 7 MIN READ
Developer Tools & Techniques

Scaling AI Inference Across Multiple GPUs Using NVIDIA TensorRT with Multi-Device Inference Support

Generative AI workloads are rapidly outgrowing the memory and compute budget of single GPUs. For inference developers building media generation pipelines, the... 11 MIN READ