Sungsoo Ha

Sungsoo Ha is a senior deep learning algorithm engineer at NVIDIA, where he works on inference optimization for LLMs. His primary technology interests are disaggregated serving, long-context decoding, and GPU parallelism strategies. He delivers NVIDIA Dynamo inference recipes and contributed decode context parallelism to vLLM and SGLang. Before NVIDIA, he optimized runtime inference for the code-editing model behind Amazon Q Developer at AWS. Sungsoo holds a PhD in Computer Science from StonyBrook University.
Avatar photo

Posts by Sungsoo Ha

Agentic AI / Generative AI

How Full-Stack NIM Optimizations Deliver 2.5x More Users on Nemotron 3 Ultra

Deploying a large language model is only the first step toward production-ready serving. Production teams also need to serve as many concurrent users as... 6 MIN READ