> Markdown version of [/jobs/ext/1714583-ai-engineer](https://www.wearedevelopers.com/jobs/ext/1714583-ai-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Engineer - **Company:** Zoom Video Communications, Inc. - **Location:** Seattle, WA, United States (Remote available) - **Experience:** Expert - **Salary:** $151,800.0 - $332,200.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, C++ (Programming Language), Nvidia CUDA, Python (Programming Language), Linux Kernel, Software Engineering, WebRTC, Large Language Models, Gpu Programming, Build Management, Information Technology, ONNX (Open Neural Network Exchange) Format, Machine Learning Operations, TensorRT, Api Design - **Published:** July 9, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=7b8706b3dc5292a0 ## About the Role * A Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a related technical field, or equivalent practical experience * 5+ years of software engineering experience, with significant time spent on inference systems or ML infrastructure at production depth * Hands-on experience with at least one major inference framework: vLLM, TensorRT-LLM, SGLang, or ONNX Runtime (serving, not just export) * GPU programming experience: CUDA kernel development, memory optimisation, and profiling with Nsight or equivalent tools * Production experience serving LLMs or large vision models - you've owned latency SLOs, debugged throughput regressions, and shipped optimisations that moved the needle * Depth in at least two of: speculative decoding, continuous batching, KV cache design, quantisation pipelines, prefill/decode disaggregation * Strong systems instincts in Python and C++; ability to read and modify framework internals Preferred * Advanced degree (Master's or PhD) in a relevant technical field * Experience with MoE models or 100B+ parameter deployments * Familiarity with disaggregated serving architectures or multi-node inference * Background in compiler-level optimisation (XLA, Triton, or similar) ## Description You'll design, implement, and own the inference systems that serve Zoom's AI models at production scale - across real-time communication, vision, and language workloads. You'll be hands-on with kernel-level optimisation, inference framework internals, and production serving infrastructure, working closely with research and platform teams to push the boundary on latency, throughput, and cost., * Design and build high-performance inference serving systems for large-scale transformer and multimodal models (including 100B+ and MoE architectures) * Implement and tune inference optimisations: speculative decoding, continuous batching, KV cache management, prefill/decode disaggregation, and quantisation (INT4/INT8/FP8) * Contribute to and customise inference frameworks (vLLM, TensorRT-LLM, SGLang, or equivalent) for Zoom's production requirements * Write and profile CUDA kernels and custom ops where framework-level optimisation is insufficient * Own end-to-end deployment: from model packaging and serving API design to latency SLO monitoring and incident response * Partner with research to translate model architecture changes into inference-efficient implementations * Drive technical design and set the bar for inference engineering practices across the team ## Related Videos - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Transforming Education: A Journey from interactive Markdown to Remote-Labs](https://www.wearedevelopers.com/videos/941-transforming-education-a-journey-from-interactive-markdown-to-remote-labs) - [API Design - Getting Started](https://www.wearedevelopers.com/videos/33-api-design-getting-started) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) - [Hello JARVIS - Building Voice Interfaces for Your LLMS](https://www.wearedevelopers.com/videos/1641-hello-jarvis-building-voice-interfaces-for-your-llms) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere)