Sai Kishan Pampana

Sai Kishan Pampana is a senior deep learning performance architect at NVIDIA, focusing on optimizing the performance of deep learning primitives on NVIDIA GPUs. His work spans both inference and training, with a particular emphasis on attention mechanisms including GQA, MLA, GDN, and other emerging variants. He has designed optimized end-to-end pipelines on NVIDIA GPUs to scale performance across real-world workloads. He holds a master's degree in Computer Science from the University of California, San Diego.
Avatar photo

Posts by Sai Kishan Pampana

Developer Tools & Techniques

Co-Designing AI Model Attention for Fast, Interactive Long-Context Inference

As agentic and long-context workloads become common, the context lengths increase and attention consumes a larger share of inference time (Figure 1). Because... 14 MIN READ