Zheng Luo

Zheng Luo is a senior software engineer on NVIDIA’s Dynamo team, where he works on distributed systems for large language model inference. His current focus is reducing LLM inference startup latency and productionizing these optimizations across widely used open-source inference frameworks. Previously, he developed and operated a large-scale inference fleet serving frontier models. Zheng holds a master’s degree from the University of California, Irvine, and a bachelor’s degree from Fudan University in Shanghai.
Avatar photo

Posts by Zheng Luo

Agentic AI / Generative AI

ModelExpress: Distributing Model Artifacts at the Speed of Light

Every byte moved has a cost. As model checkpoints grow to hundreds of gigabytes or even a terabyte, that cost adds up quickly. To make things even worse, moving... 12 MIN READ