> Markdown version of [/jobs/ext/2711234-machine-learning-engineer-inference-platform](https://www.wearedevelopers.com/jobs/ext/2711234-machine-learning-engineer-inference-platform). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Engineer (Inference Platform) - **Company:** Wizard AI, Inc. - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Microsoft Azure, Cloud Computing, Continuous Integration, DevOps, Python (Programming Language), Machine Learning, Release Management, Software Engineering, Large Language Models, Information Technology, Low Latency, Machine Learning Operations, TensorRT - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/senior-machine-learning-engineer-inference-platform-wizard-7888671 ## About the Role * Bachelor's or Master's degree in Computer Science, Data Science, Engineering, or a related field, or equivalent practical experience. * 5-8+ years of experience in Software Engineering, ML Engineering, Platform Engineering, or Infrastructure Engineering, with direct ownership of production ML serving systems. * Hands-on experience running an LLM serving engine (vLLM, TGI, TensorRT-LLM, or SGLang) in production under real load - not just managed or hosted endpoints. * Strong Python skills and software engineering fundamentals, combined with deep systems and infrastructure knowledge. * Experience with cloud platforms such as AWS, GCP, or Azure, and familiarity with ML lifecycle tooling, experimentation platforms, and model registries. * Strong grasp of inference performance - continuous batching, KV-cache and GPU-memory behavior, quantization, and CPU-versus-GPU bottlenecks - with the instinct to profile before tuning. * Experience serving heterogeneous workloads, including LLMs, embedding models, and extraction models, each with distinct latency, throughput, and scaling requirements. * Demonstrated ability to balance latency, throughput, reliability, and infrastructure cost while operating production-scale ML systems. * Experience in high-growth startup environments and comfort operating in fast-moving, evolving technical landscapes. What Success Looks Like Reliable, Scalable Inference Systems Production serving infrastructure operates with clear SLAs, strong observability, and minimal downtime. Latency, availability, throughput, and GPU utilization are actively measured and optimized as platform demands grow. ## Description As a Senior ML Engineer on our Inference Platform, you'll own the end-to-end lifecycle of production ML serving systems from model packaging and deployment to monitoring, optimization, and scaling. This is not a traditional MLOps role focused solely on pipelines and tooling. You'll be responsible for the inference infrastructure powering a live conversational shopping agent, operating multiple specialized serving engines under real-world production load. You'll own critical decisions around serving architecture, performance, reliability, and scalability, working closely with ML Engineers, Data teams, Product, and DevOps to ensure models move seamlessly from experimentation into high-performance production systems. What You'll Do * Own and evolve our multi-engine inference platform, supporting a variety of model types and serving requirements. * Build and improve production ML pipelines - taking models from experimentation to reliable, high-throughput serving. * Define and implement model versioning, rollout, rollback, and lifecycle management strategies that ensure reproducibility and operational reliability. * Define and enforce serving-layer SLAs, including latency, availability, GPU utilization, Time-to-First-Token (TTFT), and Inter-Token Latency (ITL). * Build observability, monitoring, alerting, and operational tooling for production inference systems. * Apply software engineering best practices, including testing, CI/CD integration, and reproducibility across ML workflows. * Optimize inference performance through efficient resource utilization, hardware-aware serving strategies, and cost-conscious infrastructure design. * Ensure ML serving systems are secure, scalable, and operationally resilient. * Partner with ML, Data, Product, and DevOps teams to turn ideas into production systems, driving the technical decisions on serving and scale. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Swapping Low Latency Data Storage Under High Load](https://www.wearedevelopers.com/videos/746-swapping-low-latency-data-storage-under-high-load) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) - [Unleash the power of 5G in your code: transform your apps](https://www.wearedevelopers.com/videos/1567-unleash-the-power-of-5g-in-your-code-transform-your-apps) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production)