> Markdown version of [/jobs/ext/1470829-principal-site-reliability-engineer-machine-learning](https://www.wearedevelopers.com/jobs/ext/1470829-principal-site-reliability-engineer-machine-learning). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Site Reliability Engineer, Machine Learning - **Company:** Cambridge Mobile Telematics - **Location:** Cambridge, MA, United States (Remote available) - **Experience:** Experienced - **Salary:** $142,000.0 - $177,600.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, Ubuntu (Operating System), Linux, Identity and Access Management, Python (Programming Language), Reliability Engineering, Datadog, Autoscaling, Amazon Relational Database Service, Kubernetes, Information Technology, Functional Programming, Cloudwatch, Software Coding, Amazon Simple Queue Service (SQS), Terraform, AWS EKS, Docker, Databricks, Programming Languages - **Published:** July 28, 2026 - **Apply:** https://www.dice.com/job-detail/602b1340-7e9d-4bc2-a56a-129b8bd1ea93 ## About the Role * Bachelor's degree or equivalent years of experience and/or certification in a related field * 7+ years working in Site Reliability Engineering or Information Technology * Design and document systems, including writing and reviewing code, to automate away problems within your team's domain * Intermediate to expert experience deploying and maintaining AWS services such as EC2, ECS, EKS, SQS, Lambda, Dynamo, RDS/Aurora, S3, and IAM * Intermediate to expert experience monitoring services and applications using tools such as CloudWatch Metrics, CloudWatch Logs, and Datadog, including defining and configuring alerts and SLO reports * Intermediate to expert experience maintaining the uptime and scalability of AWS compute services used for Machine Learning and Data Science workloads, specifically EC2 and EKS * Intermediate to expert coding skills in at least one programming language; we work primarily in Python * Intermediate to expert experience using Infrastructure as Code platforms and CI/CD pipelines, specifically Terraform, to manage AWS infrastructure and services * In-depth knowledge & experience with Linux operating systems (Amazon Linux, Ubuntu) on EC2 and Docker / Kubernetes * Experience with leading projects in system design, architecture changes, and technology selection ## Description * Use independent judgment and discretion to own SLOs, error budgets, and the operational health of Ray clusters running on AWS EKS and Databricks workloads on AWS EC2 across multiple accounts and regions * Maintain the observability of uptime, availability, and scalability of EKS Ray and Databricks workloads using CloudWatch and Datadog, including defining alerting that maps to SLOs * Operate and tune EKS Ray workloads at scale including autoscaling, GPU scheduling, and automated failure recovery * Manage Databricks on AWS including workspace administration, cluster policies, Unity Catalog, job orchestration, and IAM Roles and Policies * Maintain ongoing cost visibility, cost optimization, and capacity planning across EC2 and EKS workloads, including through the use of On Demand Capacity Reservations and Spot lifecycle * Perform ongoing maintenance of the underlying EC2 and EKS infrastructure, including regular security updates and operating system upgrades * Codify everything as infrastructure-as-code using Terraform and CI/CD pipelines, enabling updates through Pull Requests with approval workflows, while also automating maintenance tasks to reduce toil * Lead incident response for Data Science and Machine Learning platform outages, run blameless postmortems, and drive systemic remediation, including participating in an on-call rotation * Complete any additional tasks as they arise ## Related Videos - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Hosting a modern justice system](https://www.wearedevelopers.com/videos/332-hosting-a-modern-justice-system) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated)