Amit Kushwaha makes large language models run fast in production. As a Principal Solutions Architect at NVIDIA, he focuses on inference optimization and building agentic AI systems. Before this, he led AI engineering at SambaNova Systems and tackled physical machine learning problems at ExxonMobil.
His work gets into the practical details of low-latency inference, breaking down methods like speculative decoding and TensorRT-LLM to serve models efficiently. Amit holds a Ph.D. in Engineering from Stanford University, where his research began in scientific computing and large-scale simulations.