> Markdown version of [/jobs/ext/211618-staff-machine-learning-engineer-ml-infrastructure](https://www.wearedevelopers.com/jobs/ext/211618-staff-machine-learning-engineer-ml-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Machine Learning Engineer, ML Infrastructure - **Company:** SimpliSafe, Inc - **Location:** Boston, MA, United States - **Experience:** Expert - **Salary:** $183,500.0 - $269,100.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Amazon S3, Computer Vision, C++ (Programming Language), Cloud Computing, Code Review, Continuous Integration, Identity and Access Management, Python (Programming Language), Machine Learning, Open Source Technology, Azure Machine Learning, Video Codec, Autoscaling, Large Language Models, Containerization, Kubernetes, Apache Flink, Apache Kafka, Machine Learning Operations, TensorRT - **Published:** May 22, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=34effa313cdb2bd5 ## About the Role Do you have experience in Team leadership?, * 8+ years of software/ML engineering experience, with a clear track record of building and operating production ML systems at scale. * Deep expertise in cloud ML infrastructure on Kubernetes, with hands-on production experience with Ray (which powers our inference stack); experience with KServe, Triton, vLLM, Kubeflow, Argo, or similar is a strong plus. * Strong production experience on AWS (EKS, S3, IAM, networking) and with Kafka, containerized deployments, CI/CD, and infrastructure-as-code. * Demonstrated experience designing and operating high-throughput, low-latency inference systems - GPU-aware scheduling, batching, autoscaling, multi-tenancy. * Solid grounding in ML fundamentals: how models are trained, evaluated, versioned, deployed, monitored, and rolled back in production. * Proficiency in Python is required; experience with a systems language (Go, C++, Rust) for performance-sensitive components is a plus. * Staff-level technical leadership: ability to drive ambiguous, cross-cutting initiatives, align senior stakeholders, and elevate the engineers around you without formal authority. * Strong written and verbal communication - you can make complex technical tradeoffs legible to ML scientists, product, and other infra teams., * Hands-on experience with LLM serving in production (vLLM, TGI, TensorRT-LLM, SGLang) - KV cache management, continuous batching, speculative decoding, quantization for serving. * Experience building real-time video or streaming ML pipelines (Kafka, Kinesis, Flink, or similar) at scale. * Background supporting CV workloads in production - model formats, GPU/accelerator tradeoffs, video codecs. * Experience with model lifecycle tooling (MLflow, Weights & Biases, model registries, drift detection, shadow deployments). * Open source contributions to the ML infrastructure ecosystem (Ray, KServe, Triton, vLLM, Kubeflow, etc.). * Experience operating in environments with strong security and compliance requirements. ## Description We're looking for a Staff ML Engineer to join our Cloud ML team - the team that owns both the cloud-side ML infrastructure and the applied ML research that powers SimpliSafe's intelligent home security products. This is a senior individual contributor role focused on raising the bar for how we build, deploy, and operate ML systems at scale. You'll partner closely with other Staff and Principal engineers to drive architecture, mentor across the team, and set the technical direction for our ML platform. The work spans two of our most demanding workloads: real-time computer vision inference that processes video from cameras and doorbells across our customer base, and LLM/GenAI infrastructure that will power our future generation of intelligent applications. This role is for someone who has built ML infrastructure before, knows where the sharp edges are, and is energized by making other teams faster and more reliable. What You'll Do Set technical direction for ML infrastructure * Drive architecture decisions for our Kubernetes-based ML platform - anchored on Ray for inference, alongside KServe, Triton, and vLLM - across real-time and batch workloads. * Lead deep technical reviews on system design, capacity planning, and reliability for the highest-stakes ML systems at SimpliSafe. * Identify and remove the systemic bottlenecks in our ML deployment infrastructure - whether that's serving reliability, deployment friction, observability gaps, scaling, or cost. Build and operate real-time CV inference at scale * Own the design and evolution of cloud-side inference systems that process live video and events from SimpliSafe devices in real time. * Drive throughput, latency, and cost improvements (batching strategies, GPU utilization, autoscaling, multi-model serving) for production CV models. * Build the feedback loops between cloud inference, edge devices, and the data flywheel that improves model quality over time. Stand up LLM/GenAI serving infrastructure * Help shape how SimpliSafe serves LLMs in production - model serving patterns, KV-cache and batching strategies, evaluation pipelines, guardrails, and cost controls. * Partner with applied ML engineers to take new GenAI-powered product features from prototype to scaled deployment. Raise the engineering bar across Cloud ML * Mentor engineers across the team through design reviews, code reviews, pairing, and written guidance - a meaningful uplift on everyone you work with. * Establish and evangelize best practices for model lifecycle management (registry, deployment, monitoring, rollback, drift) and on-call. * Write the documentation, runbooks, and architectural decision records that make the platform legible and durable. Own reliability and operational excellence * Lead incident response and postmortems for critical ML systems; turn lessons learned into platform-level improvements. * Define SLOs, observability standards, and on-call practices for ML services in production., * Customer Obsessed - Building deep empathy for our customers, putting them at the core of our work, and developing strong, long-term relationships with them. * Aim High - Always challenging ourselves and others to raise the bar. * No Ego - Maintaining a "no job too small" attitude, and an open, inclusive and humble style. * One Team - Taking a highly collaborative approach to achieving success. * Lift As We Climb - Investing in developing others and helping others around us succeed. * Lean & Nimble - Working with agility and efficiency to experiment in an often ambiguous environment. ## Related Videos - [Machine Learning for Software Developers (and Knitters)](https://www.wearedevelopers.com/videos/154-machine-learning-for-software-developers-and-knitters) - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [WeAreDevelopers LIVE - CSS is DOOMed](https://www.wearedevelopers.com/videos/1838-wearedevelopers-live-css-is-doomed) - [Optimizing your AI/ML workloads for sustainability](https://www.wearedevelopers.com/videos/570-optimizing-your-ai-ml-workloads-for-sustainability) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) ## Related Articles - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How machine learning can help us tell fact from fiction](https://www.wearedevelopers.com/magazine/509-how-machine-learning-can-help-us-tell-fact-from-fiction) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer)