> Markdown version of [/jobs/ext/240635-senior-devops-sre-engineer](https://www.wearedevelopers.com/jobs/ext/240635-senior-devops-sre-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior DevOps / SRE Engineer - **Company:** Bain & Co. - **Location:** Chicago, IL, United States - **Experience:** Expert - **Salary:** $140,875.0 - $169,250.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Audit Trail, Microsoft Azure, Bash Shell, Code Review, Information Systems, Computer Networks, Continuous Integration, Data Infrastructure, DevOps, Github, Python (Programming Language), Key Management, Performance Tuning, Public Key Infrastructure, Role-Based Access Control, Reliability Engineering, Prometheus, Runbook, Tripwire, Scripting, Autoscaling, Istio, System Availability, Delivery Pipeline, Large Language Models, Grafana, Backend, Git Flow, Kubernetes, Information Technology, Hashicorp, Terraform - **Published:** May 19, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=0aab84f6561c7f90 ## About the Role * Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field (or equivalent practical experience). * 6+ years of experience in DevOps, SRE, Platform Engineering, or Production Operations roles supporting cloud-hosted, multi-service platforms. * Demonstrated experience owning production CI/CD, GitOps, and Kubernetes operations for multi-service platforms. * Experience operating and upgrading Kubernetes clusters (EKS preferred) and managing autoscaling/provisioning (e.g., Karpenter) in production. * Experience managing infrastructure-as-code at scale (Terraform), including state management and PR-driven apply workflows (e.g., Atlantis). * Track record of implementing observability and reliability practices: SLO definition, alert tuning, dashboards, incident response leadership, and post-incident reviews. * Experience operating secrets management systems (HashiCorp Vault and/or Azure Key Vault) and implementing security controls in delivery pipelines. * Strong cross-functional collaboration skills; able to enable multiple squads to deploy safely and operate services without heroics. SRE/Platform Engineering * Expert-level Kubernetes: cluster operations, upgrades, node group management (Karpenter), namespace isolation, RBAC, PodDisruptionBudgets, and topology spread. * GitOps: ArgoCD configuration, Application and Project management, sync policies, drift detection, and automated rollback patterns. * CI/CD: GitHub Actions (reusable workflows, matrix builds, secrets handling, environment protection rules, deployment gates). * Infrastructure as code: Terraform at production scale (module design, state management using Azure Storage backend, Atlantis PR-driven workflows). * Service mesh: Istio (traffic management, mTLS policy, AuthorizationPolicy, circuit breaking, observability integration). * Autoscaling and capacity: KEDA and AKS node autoprovisioning/Karpenter (event-driven autoscaling, Azure Spot VM management, bin-packing, interruption handling). * Observability: Prometheus, Grafana (dashboard-as-code), Loki, Tempo, Alertmanager (routing, inhibition, grouping); experience with Azure Monitor a plus. * Secrets management: HashiCorp Vault (auth backends, dynamic secret engines, PKI management, audit log management) and/or Azure Key Vault. * Container and supply chain security: Trivy scanning, Cosign image signing, SBOM generation, OPA/Gatekeeper policy authoring, Cilium network policy. * Scripting: strong Python and Bash for automation, tooling, and runbook automation. ## Description Senior DevOps / SRE Engineers own the CI/CD pipelines, GitOps infrastructure, Kubernetes operations, and reliability engineering practices that keep the PE platform running at production quality on Microsoft Azure. You make it safe to deploy frequently and easy to recover when things go wrong. You work closely with Platform Engineering, Data Platform, and Product squads to ensure every team can ship confidently and operate their services without heroics. WHAT YOU'LL DO Core Platform Reliability, Delivery, and Operations (80%) * Design, build, and maintain CI/CD pipelines across all repositories using reusable GitHub Actions workflows. * Own the ArgoCD GitOps configuration; manage application promotion from staging to production. * Operate and upgrade the EKS cluster; manage node groups, Karpenter provisioners, and cluster add-ons. * Maintain the Terraform estate across all environments; review and apply infrastructure changes via Atlantis. * Define and maintain SLOs, alerting rules, and Grafana dashboards for all platform services. * Operate and maintain HashiCorp Vault (and/or Azure Key Vault); manage auth backends, policies, and secret engine configuration. * Implement and maintain supply chain security controls: image scanning, signing, SBOM generation, and OPA policy enforcement. * Collaborate with the Security Engineer on network policy, egress controls, and compliance requirements. * Participate in on-call rotation; lead incident response and post-incident review process. Other (20%) * Automate repeatable operational work; reduce manual fixes through tooling and runbook automation. * Document runbooks proactively and keep them current as systems evolve. * Use AI tooling to draft infrastructure code and runbook content, validating outputs against security and compliance standards before merging. * Partner with product and engineering teams to tune reliability practices (SLOs, alerting thresholds, deployment safety checks) and to remove friction from developer workflows. * Communicate clearly during incidents: calm, factual, and action-oriented., * Integrates AI-powered quality gates into CI/CD pipelines (e.g., automated code review bots, LLM-assisted security scanning, agent-generated PR summaries for change risk assessment). * Uses AI agents to accelerate Terraform modules, Kubernetes manifests, and Helm chart scaffolding; reviews outputs against security and compliance standards before merging. * Familiar with AI-assisted incident response: using LLMs to correlate logs, suggest runbook steps, and draft post-incident reviews from structured incident data. * Contributes to Prompt Execution Sandbox and Agent Gateway infrastructure requirements from a reliability and security posture perspective. * Uses AI tooling to accelerate SLO analysis, alert rule tuning, and capacity planning modelling. General * Automates anything done more than once; prioritizes reliability and repeatability over manual fixes. * Treats SLOs as commitments, not aspirations; raises reliability concerns before they become incidents. * Documents and maintains runbooks as part of delivery, not as an afterthought. * Communicates clearly during incidents and drives structured follow-through through post-incident reviews. * This role follows a hybrid model, requiring in-office presence at least 1 day per week ## Related Videos - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Next-gen CI/CD with Gitops and Progressive Delivery](https://www.wearedevelopers.com/videos/1603-next-gen-ci-cd-with-gitops-and-progressive-delivery) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [DevOps Engineer Salary [2023]](https://www.wearedevelopers.com/magazine/203-devops-engineer-salary-2023) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [The Best X (Twitter) Accounts for Developers](https://www.wearedevelopers.com/magazine/294-the-best-x-twitter-accounts-for-developers)