Farshad Ghodsian

Farshad Ghodsian is a senior technical marketing engineer at NVIDIA, where he focuses on AI training and inference at scale, performance optimization insights, new model releases, and AI engineering enablement. He brings a wealth of experience at the intersection of AI infrastructure, distributed training, GPU-accelerated computing, and agentic workloads. With this background, he translates cutting-edge research into practical insights for developers, enterprise teams, and business leaders. Prior to NVIDIA, Farshad held technical roles at leading semiconductor and consulting companies, where he helped build and manage large-scale generative AI and MLOps platforms for top technology customers.
Avatar photo

Posts by Farshad Ghodsian

Networking / Communications

How NVIDIA NVLink 6 Delivers Multi-Layer Resiliency for AI Factories

For operators of large-scale AI factories, maximizing continuous output is essential for productivity. In massive-scale AI training, every GPU in the cluster... 12 MIN READ
Data Center / Cloud

NVIDIA NVLink Fusion Brings NVHBM to Next-Generation AI Infrastructure

AI factories must support increasingly large models and more complex reasoning workloads. To keep up with the insatiable compute demands of AI workloads,... 7 MIN READ
Agentic AI / Generative AI

Setting a World Record for MoE Pre-Training on NVIDIA GB300 NVL72

Frontier model pre-training has converged on mixture of experts (MoE), which is fundamentally changing what limits large-scale AI training. As compute per token... 8 MIN READ
Data Center / Cloud

Inside NVIDIA Rubin GPU Architecture: Powering the Era of Agentic AI

What began as discrete AI model training and human-facing chat interfaces has evolved into always-on AI factories dedicated to producing intelligence at scale.... 15 MIN READ
Agentic AI / Generative AI

NVIDIA Vera CPU Boosts AI Factory Throughput to Accelerate Agentic Workloads

Agentic systems turn model reasoning into action through multi-step workflows that combine inference, tool use, code execution, retrieval, orchestration, and... 8 MIN READ
Data Center / Cloud

Enhancing Goodput in Large-Scale LLM Training with Nonuniform Tensor Parallelism

Training LLMs at massive scale brings unique infrastructure challenges, especially as jobs span thousands of GPUs and run for extended periods. The longer these... 7 MIN READ