Kavin Krishnan

Kavin Krishnan is a deep learning systems software engineer at NVIDIA, where he has spent 6 years accelerating distributed inference at scale. As a member of the founding team of ModelExpress, he helped develop a fast P2P solution for reducing cold-start latency for distributed LLM workloads. His current work focuses on optimizing weight refits between trainers and generators in post-training reinforcement learning workflows. He holds a master’s degree in machine learning from Georgia Tech.
Avatar photo

Posts by Kavin Krishnan

Agentic AI / Generative AI

ModelExpress: Distributing Model Artifacts at the Speed of Light

Every byte moved has a cost. As model checkpoints grow to hundreds of gigabytes or even a terabyte, that cost adds up quickly. To make things even worse, moving... 12 MIN READ