> Markdown version of [/jobs/ext/2522553-senior-devops-mlops-engineer](https://www.wearedevelopers.com/jobs/ext/2522553-senior-devops-mlops-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior DevOps / MLOps Engineer - **Company:** Simpligov Llc - **Location:** Baltimore, MD, United States (Remote available) - **Experience:** Expert - **Salary:** $160,000.0 - $190,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Audit Trail, Microsoft Azure, Software as a Service, DevOps, Key Management, Octopus Deploy, Role-Based Access Control, Reliability Engineering, AI Infrastructure, Large Language Models, Snowflake, Kubernetes, Bicep, Machine Learning Operations, Terraform, Microservices - **Published:** August 11, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=c0d9aca34d3abf32 ## About the Role * 5+ years in DevOps, platform engineering, or site reliability engineering in SaaS environments * Deep Azure experience: AKS, networking, identity (Entra), and monitoring; you have run production Kubernetes * Infrastructure as code as your default (Terraform, Bicep, or similar), plus strong scripting; you automate before you document * MLOps experience: deploying and operating LLM or ML systems in production, including model gateways, inference infrastructure, or AI observability stacks * Demonstrated cost work: you can point to cloud spend you found, explained, and reduced * Experience in compliance-heavy environments (FedRAMP, StateRAMP, SOC 2, or similar) is a strong plus * Comfortable holding production access, with the discipline that implies Key Competencies * Treats environment integrity as sacred: no invisible changes, no snowflake servers, no heroics that cannot be audited * Cost literacy: reads a cloud bill the way an engineer reads a stack trace * Automates first: your instinct is a pipeline or a policy, not a runbook step * Thinks in the open: surfaces risk early and documents what you build * Calm in production incidents; rigorous in the postmortem ## Description You will own the Azure platform behind SimpliGov's AI-native delivery model-infrastructure, Kubernetes, networking, observability, AI serving, and cost discipline. This is hands-on production engineering within a FedRAMP-conscious environment, where security, auditability, and reliability are core responsibilities. Responsibilities: * Deploy and operate our Azure platform: AKS, networking, identity, storage, and environments from development through production * Own infrastructure as code end to end: environments are reproducible, drift is detected, and nothing reaches an environment without platform visibility * Operate the AI infrastructure layer: self-hosted observability and evaluation tooling (Langfuse), product telemetry, model gateway and per-workload routing, and compliant GovCloud inference paths * Own cloud and AI cost: metering, budgets, unit economics, MACC drawdown strategy, and active remediation; cost is an engineering metric here, not a finance afterthought * Harden production access and controls: least privilege, secrets management, audit evidence, and a FedRAMP-conscious security posture * Partner with AI Operations on the deploy-and-release path: Octopus Deploy, environment promotion, progressive rollout, and rollback * Build platform reliability: monitoring, alerting, incident response, and capacity planning * Give the microservices decomposition the platform primitives it needs: service infrastructure, scaling patterns, and clean environment boundaries ## Related Videos - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [From Traction to Production: Maturing your LLMOps step by step](https://www.wearedevelopers.com/videos/1250-from-traction-to-production-maturing-your-llmops-step-by-step) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [Back(end) to the Future: Embracing the continuous Evolution of Infrastructure and Code](https://www.wearedevelopers.com/videos/440-back-end-to-the-future-embracing-the-continuous-evolution-of-infrastructure-and-code) - [Implementing Feature Environments with AWS and Terraform](https://www.wearedevelopers.com/videos/531-implementing-feature-environments-with-aws-and-terraform) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)