Senior Site Reliability Engineer (SRE)

Concentrix Corporation
Bellevue, WA, United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$92,250.0 - $140,000.0
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Agile Methodology Artificial Intelligence Amazon Web Services Apple IOS Microsoft Azure Bash Shell Cloud Computing Cloud Engineering Configuration Management Continuous Integration DevOps
+30 more
Disaster Recovery Distributed Systems Monitoring of Systems Mobile Application Software Python (Programming Language) Linux System Administration Reliability Engineering Prometheus Azure Machine Learning Software Engineering Web Platforms Datadog Cloud Platform System System Availability Grafana Mttr Git Cloudformation Containerization Kubernetes Information Technology Deployment Automation Api Design Api Gateway Terraform Splunk New Relic (SaaS) Devsecops Docker Microservices

Job description

As a Senior Site Reliability Engineer , you will help build, scale, and operate the resilient platforms that power critical digital experiences across web, mobile, API, AI/ML, and customer-facing environments. This role is ideal for an engineer who thrives at the intersection of software, cloud infrastructure, automation, and operations, and who is passionate about improving reliability, observability, scalability, and developer experience.

You will collaborate closely with software engineering, platform, architecture, product, and security teams to strengthen platform performance and availability, reduce operational toil, and advance intelligent, automated operations across mission-critical services.

Responsibilities

  • Design, implement, and support highly available, scalable, and resilient platform services.

  • Define and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets.

  • Identify and address reliability risks and performance bottlenecks across distributed systems.

  • Lead root cause analysis for critical incidents and drive sustainable corrective actions.

  • Participate in production support and on-call rotations for business-critical applications.

  • Build and enhance observability solutions using tools such as Splunk, Grafana, Prometheus, Datadog, New Relic, and OpenTelemetry.

  • Create actionable dashboards, alerts, metrics, and reporting that improve operational visibility.

  • Drive continuous improvement in Mean Time to Detect (MTTD) and Mean Time to Resolution (MTTR).

  • Support and evolve shared platform capabilities used across multiple engineering teams.

  • Develop self-service platform features, reusable services, and automation frameworks that improve developer productivity.

  • Partner with architecture and engineering teams to define future-state platform strategies.

  • Design and manage cloud-native infrastructure across AWS and Azure environments.

  • Implement Infrastructure as Code using Terraform, CloudFormation, Helm, and Kubernetes manifests.

  • Automate provisioning, deployment, configuration management, and recovery processes.

  • Design and optimize CI/CD pipelines to enable secure, reliable, and efficient software delivery.

  • Improve deployment speed and quality through automation, release validation, and deployment controls.

  • Support GitOps operating models and deployment automation practices.

  • Lead operational readiness reviews, disaster recovery exercises, and resiliency initiatives.

  • Establish runbooks, playbooks, and automated remediation solutions.

  • Drive chaos engineering and resiliency testing efforts.

  • Ensure alignment with enterprise operational standards and security requirements.

  • Leverage AI-assisted operations, observability, and incident intelligence capabilities to proactively identify and mitigate risk.

  • Advance intelligent platform capabilities that enhance engineering efficiency and operational excellence.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related field, or equivalent practical experience.

  • 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Cloud Operations.

  • Experience supporting large-scale distributed applications in production environments.

  • Experience operating customer-facing digital platforms with high availability expectations.

  • Strong Linux administration and troubleshooting skills.

  • Expertise with Kubernetes and containerized environments.

  • Experience with AWS and/or Azure cloud platforms.

  • Strong scripting and automation skills in Python and Bash; Go is preferred.

  • Strong experience with Terraform, Git, CI/CD pipelines, Docker, and Kubernetes.

  • Experience working with APIs, microservices, and distributed architectures.

  • Hands-on experience with observability and monitoring tools such as Splunk, Grafana, Prometheus, OpenTelemetry, New Relic, and cloud-native monitoring platforms.

  • Demonstrated strength in incident management, escalation leadership, root cause analysis, and problem management.

  • Experience with capacity planning, performance engineering, disaster recovery, and resiliency testing.

  • Preferred: experience supporting mobile applications (iOS and Android) and digital customer platforms.

  • Preferred: experience with API gateways, CDN technologies, and edge architectures.

  • Preferred: knowledge of AI/ML platform operations.

  • Preferred: familiarity with cybersecurity best practices and DevSecOps.

  • Preferred: AWS Certified DevOps Engineer, Solutions Architect, Kubernetes, or related certifications.

  • Preferred: experience working within Agile and Product operating models., In accordance with federal law, only applicants who are legally authorized to work in the United States will be considered for this position. Must reside in the United States or have a valid U.S. address for residence.

Benefits & conditions

The base salary range for this position is $92,250 - $140,000 plus incentives that align with individual and company performance. Actual salaries will vary based on work location, qualifications, skills, education, experience, and competencies. Benefits available to eligible employees in this role include medical, dental, and vision insurance, comprehensive employee assistance program, 401(k) retirement plan, paid time off and holidays.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.juju.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

3:08 min

Aligning engineering processes with core business impact metrics

Chris Riley · WWC 2021

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · WWC 2023

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

Videos

See all

Related articles

See all