Principal Site Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+11 more
Job description
Experteer Overview In this role you will strengthen reliability across cloud and on-prem environments by applying SRE principles, automation, and observability. You’ll lead chaos testing initiatives, build scalable automation, and mentor teams to achieve resilient platform availability. You’ll partner with product and platform groups to embed reliability into roadmaps and practices, driving measurable improvements in system stability. This is an on-site heavy, technically focused opportunity at Fidelity, with impact across workplace investing, healthcare, and defined benefits domains. Compensation / Benefits * Provide cloud support and improve cloud capabilities following SRE principles (observability, automation, resiliency) * Develop and enhance internal chaos framework for chaos executions and reporting * Facilitate chaos engineering adoption by application teams; conduct chaos testing and analyze weaknesses to boost resiliency * Design and develop products within the SRE domain to improve stability and platform availability * Collaborate with business and technology teams to scale products and automation across units * Develop strategies and tools to remediate operational problems and minimize impact * Offer technical leadership on chaos testing for cloud and on-premises applications * Create scripts and applications to automate repeatable business processes * Advise senior management on technical strategy and tooling * Mentor team members to build core SRE competencies Tasks * Bachelor’s degree in Computer Science, Engineering, Information Technology, Information Systems, or closely related field with five years of experience as a Principal SRE or equivalent * Or Master’s degree with three years of experience as a Principal SRE or equivalent * Proven experience designing and automating container and cloud-based platform products in production environments * Strong knowledge of Kubernetes and containerized workloads * Experience with infrastructure-as-code tools (Azure ARM, Terraform) and cloud platforms (AWS, Azure) * Experience with monitoring, logging, and alerting of distributed systems (Datadog, Splunk) * Proficiency with DevOps tools (Jenkins, Azure DevOps, Team Foundation Version Control, CloudFormation) * Experience with AWS Lambda, API Gateway, FIS, and Azure Chaos Studio; familiarity with Windows and Linux scripting (Python) * Ability to develop chaos testing frameworks and drive adoption across teams * Strong leadership, mentoring, and stakeholder collaboration skills Key requirements *
Requirements
Advise stability and platform availability * Collaborate with business and technology teams to scale products and automation across units * Develop strategies and tools to remediate operational problems and minimize impact * Offer technical leadership on chaos testing for cloud and on-premises applications * Create scripts and applications to automate repeatable business processes * Advise senior management on technical strategy and tooling * Mentor team members to build core SRE competencies Tasks * Bachelor’s degree in Computer Science, Engineering, Information Technology, Information Systems, or closely related field with five years of experience as a Principal SRE or equivalent * Or Master’s degree with three years of experience as a Principal SRE or equivalent * Proven experience designing and automating container and cloud-based platform products in production environments * Strong knowledge of Kubernetes and containerized workloads * Experience with infrastructure-as-code tools a tools ARM, Terraform) and cloud platforms (AWS, Azure) * Experience with monitoring, logging, and alerting of distributed systems (Datadog, Splunk) * Proficiency with DevOps tools (Jenkins, Azure DevOps, Team Foundation Version Control, CloudFormation) * Experience with AWS Lambda, API Gateway, FIS, and Azure Chaos Studio; familiarity with Windows and Linux scripting (Python) * Ability to develop chaos testing frameworks and drive adoption across teams * Strong leadership, mentoring, and stakeholder collaboration skills Key requirements *
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
What Are The Top Skills Required For Azure Developers?
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
Is Software Engineering Over-Saturated?
Why Upskilling And Reskilling is Important For Developers