> Markdown version of [/jobs/ext/938696-sr-sre-dev-ops-engineer](https://www.wearedevelopers.com/jobs/ext/938696-sr-sre-dev-ops-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Sr SRE/Dev Ops Engineer - **Company:** THE MADISON - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $170,000.0 - $175,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Build Automation, Automation of Tests, Bash Shell, Cloud Computing, Computer Programming, System Configuration, Continuous Integration, DevOps, Identity and Access Management, Python (Programming Language), Operational Data Store, Queueing Systems, Reliability Engineering, Prometheus, Data Streaming, Datadog, Data Logging, Scripting, Cloud Platform System, Delivery Pipeline, Grafana, Reliability of Systems, Cloudformation, Event Driven Architecture, AI Platforms, Kubernetes, Infrastructure Automation Frameworks, Deployment Automation, Data Analytics, Machine Learning Operations, Terraform, Splunk, New Relic (SaaS), Software Version Control, Golang - **Published:** June 30, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=e79f31094a9f2ff9 ## About the Role Do you have experience in Tooling?, Madison Reed is seeking a hands-on Senior SRE / AI Platform DevOps Engineer to build, operate, and scale the infrastructure behind our AI-powered services, agents, and orchestration platforms. This role sits at the intersection of site reliability engineering, cloud infrastructure, DevOps automation, observability, and AI operations. You will own the systems and practices that ensure our AI-enabled services are reliable, secure, scalable, cost-effective, and production-ready. The ideal candidate is infrastructure-first and operationally minded, with deep experience in cloud environments, CI/CD, production monitoring, incident response, and automation. You will help operationalize AI systems by building reliable deployment workflows, telemetry pipelines, monitoring frameworks, and governance processes for models, agents, and orchestration services. This is a highly hands-on engineering role for someone who enjoys building resilient platforms, reducing operational risk, improving deployment velocity, and making advanced technology dependable in real-world production environments., * 5+ years of experience in DevOps, Site Reliability Engineering, Platform Engineering, Cloud Infrastructure, or related roles. * Strong hands-on experience with cloud infrastructure, preferably AWS. * Experience building and maintaining CI/CD pipelines and automated deployment workflows. * Proficiency with infrastructure-as-code tools such as Terraform, CloudFormation or similar. * Experience operating production systems with strong monitoring, alerting, logging, and incident response practices. * Strong scripting or programming skills in Python, Bash, Go, or similar languages. * Experience designing reliable, secure, scalable, and cost-conscious infrastructure. * Comfortable participating in on-call rotations and supporting production systems., * Experience operating AI, ML, agent-based, or data-intensive systems in production. * Familiarity with model deployment, model versioning, inference services, or MLOps workflows. * Experience with observability platforms such as Datadog, New Relic, Grafana, Prometheus, OpenTelemetry, Splunk, or similar. * Experience with event-driven architectures, queueing systems, streaming platforms, or telemetry pipelines. * Familiarity with AIOps concepts such as anomaly detection, alert correlation, automated remediation, and intelligent incident response. * Experience implementing SLOs, error budgets, production readiness reviews, and reliability scorecards. * Understanding of security, compliance, access control, and governance practices for production systems. ## Description * Design, provision, and manage cloud infrastructure for AI-powered services, agents, orchestration systems, and supporting platforms. * Automate environment setup and configuration across development, staging, and production environments. * Build reusable infrastructure-as-code patterns that improve consistency, security, scalability, and maintainability. * Partner with engineering teams to ensure production systems are resilient, observable, performant, and cost-efficient. * Participate in on-call support, incident response, root cause analysis, and continuous reliability improvement., * Build, maintain, and optimize CI/CD pipelines for services, agents, orchestration layers, and supporting infrastructure. * Implement automated testing, validation, security, and reliability gates within deployment workflows. * Design safe deployment patterns including blue/green deployments, canary releases, feature flags, and automated rollback mechanisms. * Integrate health checks, service readiness checks, and reliability signals into release processes. * Improve deployment speed and confidence while reducing production risk., * Package, version, deploy, and manage AI models, agent services, and orchestration components across environments. * Support safe rollout, rollback, refresh, and retirement workflows for AI-powered services. * Monitor AI service performance across latency, throughput, availability, cost, quality, and business-critical reliability signals. * Implement operational controls for AI systems, including version tracking, environment promotion, access management, and change governance. * Partner with data, engineering, product, and support teams to ensure AI systems are production-ready and operationally accountable., * Design and operate scalable telemetry pipelines for logs, metrics, traces, model events, agent interactions, and operational signals. * Enable structured observability for AI services and orchestration systems to support real-time monitoring, alerting, and diagnostics. * Build dashboards, alerts, and reporting that provide actionable insight into system health, performance, reliability, and cost. * Improve incident detection, triage, and resolution through high-quality telemetry and operational data. * Support data-driven reliability practices, including SLOs, error budgets, service health reviews, and post-incident analysis., * Implement intelligent monitoring, alert correlation, anomaly detection, and automated incident response capabilities. * Integrate AIOps tools and workflows into existing DevOps, SRE, and engineering operations. * Build automation that reduces manual operational work and improves mean time to detect and resolve issues. * Identify opportunities to use AI and automation to improve platform reliability, observability, supportability, and operational efficiency. Production Reliability & SRE Excellence * Define and maintain reliability standards for AI-powered production systems. * Establish and track service-level indicators, service-level objectives, and operational readiness requirements. * Lead reliability reviews, production readiness assessments, and infrastructure risk assessments. * Drive improvements in system resilience, scalability, security, performance, and cost optimization. * Champion SRE best practices across engineering teams. ## Related Videos - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [What is Software Engineering in the Age of AI?](https://www.wearedevelopers.com/magazine/640-what-is-software-engineering-in-the-age-of-ai) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)