Philip Kiely scales AI models as part of the special projects team at Baseten. He wrote the book on Inference Engineering, focusing on the complex puzzle of latency, throughput, and cost in generative AI.
His technical work spans the serving stack from CUDA to Kubernetes. He tackles the hard problems of topology-aware model parallelism and infrastructure management required to run massive LLMs efficiently in production.
Based in San Francisco, Philip spends his time practicing martial arts, reading, or cheering for his adopted local sports teams.