Site Reliability Engineer IV
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+13 more
Job description
Experteer Overview In this role you will lead platform reliability across the enterprise, acting as a SME in Site Reliability Engineering. You will drive reliability standards, observability, automation, and incident response to improve system stability and performance. You will mentor engineers and influence enterprise engineering practices while partnering with senior stakeholders. This is a scale-focused position at a financial services firm with a strong emphasis on risk management and operational excellence. Compensation / Benefits * Define and drive service reliability standards (SLOs, SLAs, SLIs, error budgets) * Design highly available, fault-tolerant architectures * Lead automation to improve reliability and operational excellence * Develop observability strategy using logging, monitoring, tracing, dashboards, and telemetry analytics * Design and maintain end-to-end monitoring solutions for application, infra, and customer experience * Analyze production telemetry to identify performance bottlenecks and risks * Lead incident management and post-incident RCA activities * Drive automation for self-healing systems and deployment/recovery workflows * Partner with development teams to build observable, scalable services across SDLC * Develop automated regression testing strategies and validate stability and performance * Create and improve IaC solutions using Terraform * Support Azure cloud environments and deployment automation * Engage in performance engineering, resilience, capacity planning, and workload optimization * Lead production readiness activities and DR/operational readiness reviews * Review architectures and roadmaps for reliability improvements * Mentor engineers on reliability, observability, cloud engineering, and automation * Prepare operational runbooks, incident playbooks, and knowledge articles * Communicate reliability metrics and remediation strategies to stakeholders * Participate in architecture reviews and leadership discussions * Promote a culture of belonging and compliance with internal controls and regulatory requirements * Complete other related duties as assigned Tasks * Associate’s degree with 9+ years of experience or Bachelor’s degree with 7+ years of experience in systems analysis and/or application development * Expert experience in system design, reliability engineering, and production operations * Advanced proficiency in at least one programming or scripting language * Experience with observability and incident management tooling * Experience with cloud platforms (AWS or Azure) * Strong understanding of CI/CD, DevOps, and SDLC practices * Experience defining and implementing SLO/SLI frameworks * Experience in regulated environments such as financial services * Experience with Infrastructure as Code (Terraform) * Experience with monitoring/observability tools (Dynatrace, OpenTelemetry, Azure Monitor, Application Insights) * Experience with automated testing, deployment automation, and reliability engineering practices * Knowledge of capacity planning, resiliency testing, disaster recovery, and high-availability architectures * Experience with Agile/DevOps operating models * Scripting/automation in PowerShell, Python, Bash (or similar) * Industry cloud certifications preferred Key requirements *
Requirements
Promote a culture of belonging and compliance with internal controls and regulatory requirements * Complete other related duties as assigned Tasks * Associate’s degree with 9+ years of experience or Bachelor’s degree with 7+ years of experience in systems analysis and/or application development * Expert experience in system design, reliability engineering, and production operations * Advanced proficiency in at least one programming or scripting language * Experience with observability and incident management tooling * Experience with cloud platforms (AWS or Azure) * Strong understanding of CI/CD, DevOps, and SDLC practices * Experience defining and implementing SLO/SLI frameworks * Experience in regulated environments such as financial services * Experience with Infrastructure as Code (Terraform) * Experience with monitoring/observability tools (Dynatrace, OpenTelemetry, Azure Monitor, Application Insights) * Experience with automated testing, deployment automation, and aaaaaa aaaN_ engineering practices * Knowledge of capacity planning, resiliency testing, disaster recovery, and high-availability architectures * Experience with Agile/DevOps operating models * Scripting/automation in PowerShell, Python, Bash (or similar) * Industry cloud certifications preferred Key requirements *
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on us.experteer.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Best Paying Jobs in Technology
The Most Popular IT Jobs on the Market
Top-Paying Tech Jobs (with Salaries)
Highest Paying Tech Companies for Developers