Principal Site Reliability Engineer

Fmr LLC
Westlake, LA, United States
6 days ago
Apply on us.experteer.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours

Tech stack

Amazon Web Services Microsoft Azure Cloud Engineering Information Systems DevOps Distributed Systems Python (Programming Language) Reliability Engineering Cloud Services Shell Script Windows Desktop Datadog
+11 more
Data Logging Cloudformation Containerization Kubernetes Infrastructure Automation Frameworks Information Technology Api Gateway Terraform Splunk Software Version Control Jenkins

Job description

Experteer Overview In this role you will strengthen reliability across cloud and on-prem environments by applying SRE principles, automation, and observability. You’ll lead chaos testing initiatives, build scalable automation, and mentor teams to achieve resilient platform availability. You’ll partner with product and platform groups to embed reliability into roadmaps and practices, driving measurable improvements in system stability. This is an on-site heavy, technically focused opportunity at Fidelity, with impact across workplace investing, healthcare, and defined benefits domains. Compensation / Benefits * Provide cloud support and improve cloud capabilities following SRE principles (observability, automation, resiliency) * Develop and enhance internal chaos framework for chaos executions and reporting * Facilitate chaos engineering adoption by application teams; conduct chaos testing and analyze weaknesses to boost resiliency * Design and develop products within the SRE domain to improve stability and platform availability * Collaborate with business and technology teams to scale products and automation across units * Develop strategies and tools to remediate operational problems and minimize impact * Offer technical leadership on chaos testing for cloud and on-premises applications * Create scripts and applications to automate repeatable business processes * Advise senior management on technical strategy and tooling * Mentor team members to build core SRE competencies Tasks * Bachelor’s degree in Computer Science, Engineering, Information Technology, Information Systems, or closely related field with five years of experience as a Principal SRE or equivalent * Or Master’s degree with three years of experience as a Principal SRE or equivalent * Proven experience designing and automating container and cloud-based platform products in production environments * Strong knowledge of Kubernetes and containerized workloads * Experience with infrastructure-as-code tools (Azure ARM, Terraform) and cloud platforms (AWS, Azure) * Experience with monitoring, logging, and alerting of distributed systems (Datadog, Splunk) * Proficiency with DevOps tools (Jenkins, Azure DevOps, Team Foundation Version Control, CloudFormation) * Experience with AWS Lambda, API Gateway, FIS, and Azure Chaos Studio; familiarity with Windows and Linux scripting (Python) * Ability to develop chaos testing frameworks and drive adoption across teams * Strong leadership, mentoring, and stakeholder collaboration skills Key requirements *

Requirements

Advise stability and platform availability * Collaborate with business and technology teams to scale products and automation across units * Develop strategies and tools to remediate operational problems and minimize impact * Offer technical leadership on chaos testing for cloud and on-premises applications * Create scripts and applications to automate repeatable business processes * Advise senior management on technical strategy and tooling * Mentor team members to build core SRE competencies Tasks * Bachelor’s degree in Computer Science, Engineering, Information Technology, Information Systems, or closely related field with five years of experience as a Principal SRE or equivalent * Or Master’s degree with three years of experience as a Principal SRE or equivalent * Proven experience designing and automating container and cloud-based platform products in production environments * Strong knowledge of Kubernetes and containerized workloads * Experience with infrastructure-as-code tools a tools ARM, Terraform) and cloud platforms (AWS, Azure) * Experience with monitoring, logging, and alerting of distributed systems (Datadog, Splunk) * Proficiency with DevOps tools (Jenkins, Azure DevOps, Team Foundation Version Control, CloudFormation) * Experience with AWS Lambda, API Gateway, FIS, and Azure Chaos Studio; familiarity with Windows and Linux scripting (Python) * Ability to develop chaos testing frameworks and drive adoption across teams * Strong leadership, mentoring, and stakeholder collaboration skills Key requirements *

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:02 min

Applying an ETL methodology to infrastructure configuration management

Axel Barbier · World Congress 2023

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

57 sec

Extracting API schemas automatically during continuous integration builds

Axel Barbier · World Congress 2023

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all