> Markdown version of [/jobs/ext/3545641-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/3545641-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Capgemini - **Location:** Boston, MA, United States - **Experience:** Expert - **Salary:** $86,129.0 - $127,189.0 - **Contract:** Permanent contract - **Skills:** Microsoft Windows, Amazon Web Services, Microsoft Azure, Big Data, Continuous Integration, Query Languages, Linux, DevOps, Distributed Systems, Python (Programming Language), Node.Js, Reliability Engineering, Power BI, Software Deployment, Software Engineering, Software Organization, Datadog, Data Logging, Scripting, Delivery Pipeline, Grafana, Reliability of Systems, Cloudformation, Kubernetes, Infrastructure Automation Frameworks, Deployment Automation, Hardware Infrastructure, Api Gateway, Terraform, Splunk, Software Version Control, Jenkins - **Published:** October 1, 2026 - **Apply:** https://www.capgemini.com/jobs/568622-en_US_SAPBTP/x/ ## About the Role * 9+ years of demonstrated experience developing and designing products around Site Reliability Engineering principles to improve stability and platform availability for containerized workloads and on-premises services using Kubernetes. * Experience managing and interpreting large datasets using query languages and creating dashboards and reports with Power BI and Grafana. * Strong background in managing cloud and on-premises infrastructure using Infrastructure as Code tools, including Terraform, and CloudFormation. * Hands-on experience building, operating, monitoring, logging, and alerting distributed systems at scale using Datadog and Splunk. * Experience supporting DevOps practices for service delivery and operations using Jenkins, Azure DevOps, Team Foundation Version Control, and CI/CD automation. * Experience developing software and automation solutions to support application delivery, operations, and repeatable business processes using Python. * Knowledge of scalability and resiliency practices for applications deployed on AWS and Azure, including Lambda and API Gateway * Strong development experience in scripting, automation, and integration across Linux and Windows-based environments. ## Description We are seeking a highly motivated Site Reliability Engineer to help build and operate reliable, scalable, and secure services across our platform. This role is designed for someone who combines strong DevOps practices, modern SRE principles, and software engineering experience to improve system reliability, automate operations, and support high-availability production environments. The ideal candidate will be passionate about building resilient systems, improving developer productivity, and driving operational excellence through automation, observability, and engineering best practices. This role partners closely with engineering, platform, and product teams to ensure services are built for reliability from the start and remain performant, stable, and supportable at scale Key Skills - Node.js, Python, DevOps, Jenkins, AWS, * Design, build and operate resilient, scalable systems using DevOps, SRE, and software development best practices. * Deliver high-availability services through automation, infrastructure as code, and proactive reliability engineering. * Improve monitoring, logging, alerting, and observability for distributed systems. * Support CI/CD automation, deployment workflows, and production tooling to reduce operational toil. * Drive incident response, root cause analysis, and recovery improvements to minimize downtime. * Partner with engineering teams to embed reliability into the software development lifecycle. * Automate provisioning, configuration, and self-healing across cloud and on-prem environments. * Validate resiliency and performance through testing, chaos engineering, and capacity planning.' ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Shifting Stress to Progress— Understanding DevOps to do DevOps Better](https://www.wearedevelopers.com/videos/268-shifting-stress-to-progress-understanding-devops-to-do-devops-better) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) ## Related Articles - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)