> Markdown version of [/jobs/ext/2034566-principal-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2034566-principal-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Site Reliability Engineer - **Company:** Fmr LLC - **Location:** Westlake, OH, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Microsoft Azure, Cloud Engineering, Information Systems, DevOps, Distributed Systems, Python (Programming Language), Reliability Engineering, Cloud Services, Shell Script, Windows Desktop, Datadog, Data Logging, Cloudformation, Containerization, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Api Gateway, Terraform, Splunk, Software Version Control, Jenkins - **Published:** August 12, 2026 - **Apply:** https://us.experteer.com/career/view-jobs/principal-site-reliability-engineer-westlake-oh-usa-58918037 ## About the Role Advise stability and platform availability * Collaborate with business and technology teams to scale products and automation across units * Develop strategies and tools to remediate operational problems and minimize impact * Offer technical leadership on chaos testing for cloud and on-premises applications * Create scripts and applications to automate repeatable business processes * Advise senior management on technical strategy and tooling * Mentor team members to build core SRE competencies Tasks * Bachelor's degree in Computer Science, Engineering, Information Technology, Information Systems, or closely related field with five years of experience as a Principal SRE or equivalent * Or Master's degree with three years of experience as a Principal SRE or equivalent * Proven experience designing and automating container and cloud-based platform products in production environments * Strong knowledge of Kubernetes and containerized workloads * Experience with infrastructure-as-code tools a tools ARM, Terraform) and cloud platforms (AWS, Azure) * Experience with monitoring, logging, and alerting of distributed systems (Datadog, Splunk) * Proficiency with DevOps tools (Jenkins, Azure DevOps, Team Foundation Version Control, CloudFormation) * Experience with AWS Lambda, API Gateway, FIS, and Azure Chaos Studio; familiarity with Windows and Linux scripting (Python) * Ability to develop chaos testing frameworks and drive adoption across teams * Strong leadership, mentoring, and stakeholder collaboration skills Key requirements * ## Description Experteer Overview In this role you will strengthen reliability across cloud and on-prem environments by applying SRE principles, automation, and observability. You'll lead chaos testing initiatives, build scalable automation, and mentor teams to achieve resilient platform availability. You'll partner with product and platform groups to embed reliability into roadmaps and practices, driving measurable improvements in system stability. This is an on-site heavy, technically focused opportunity at Fidelity, with impact across workplace investing, healthcare, and defined benefits domains. Compensation / Benefits * Provide cloud support and improve cloud capabilities following SRE principles (observability, automation, resiliency) * Develop and enhance internal chaos framework for chaos executions and reporting * Facilitate chaos engineering adoption by application teams; conduct chaos testing and analyze weaknesses to boost resiliency * Design and develop products within the SRE domain to improve stability and platform availability * Collaborate with business and technology teams to scale products and automation across units * Develop strategies and tools to remediate operational problems and minimize impact * Offer technical leadership on chaos testing for cloud and on-premises applications * Create scripts and applications to automate repeatable business processes * Advise senior management on technical strategy and tooling * Mentor team members to build core SRE competencies Tasks * Bachelor's degree in Computer Science, Engineering, Information Technology, Information Systems, or closely related field with five years of experience as a Principal SRE or equivalent * Or Master's degree with three years of experience as a Principal SRE or equivalent * Proven experience designing and automating container and cloud-based platform products in production environments * Strong knowledge of Kubernetes and containerized workloads * Experience with infrastructure-as-code tools (Azure ARM, Terraform) and cloud platforms (AWS, Azure) * Experience with monitoring, logging, and alerting of distributed systems (Datadog, Splunk) * Proficiency with DevOps tools (Jenkins, Azure DevOps, Team Foundation Version Control, CloudFormation) * Experience with AWS Lambda, API Gateway, FIS, and Azure Chaos Studio; familiarity with Windows and Linux scripting (Python) * Ability to develop chaos testing frameworks and drive adoption across teams * Strong leadership, mentoring, and stakeholder collaboration skills Key requirements * ## Related Videos - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [The Road to MLOps: How Verivox Transitioned to AWS](https://www.wearedevelopers.com/videos/1050-the-road-to-mlops-how-verivox-transitioned-to-aws) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Our GitOps approach for deploying an Identity Provider and an API Gateway in a SaaS company](https://www.wearedevelopers.com/videos/776-our-gitops-approach-for-deploying-an-identity-provider-and-an-api-gateway-in-a-saas-company) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)