> Markdown version of [/jobs/ext/1768703-senior-site-reliability-engineer-sre](https://www.wearedevelopers.com/jobs/ext/1768703-senior-site-reliability-engineer-sre). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer (SRE) - **Company:** Concentrix Corporation - **Location:** Bellevue, WA, United States - **Experience:** Expert - **Salary:** $92,250.0 - $140,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Agile Methodology, Artificial Intelligence, Amazon Web Services, Apple IOS, Microsoft Azure, Bash Shell, Cloud Computing, Cloud Engineering, Configuration Management, Continuous Integration, DevOps, Disaster Recovery, Distributed Systems, Monitoring of Systems, Mobile Application Software, Python (Programming Language), Linux System Administration, Reliability Engineering, Prometheus, Azure Machine Learning, Software Engineering, Web Platforms, Datadog, Cloud Platform System, System Availability, Grafana, Mttr, Git, Cloudformation, Containerization, Kubernetes, Information Technology, Deployment Automation, Api Design, Api Gateway, Terraform, Splunk, New Relic (SaaS), Devsecops, Docker, Microservices - **Published:** July 11, 2026 - **Apply:** https://www.juju.com/job/00000000gfke1q ## About the Role + Bachelor's degree in Computer Science, Engineering, Information Technology, or a related field, or equivalent practical experience. + 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Cloud Operations. + Experience supporting large-scale distributed applications in production environments. + Experience operating customer-facing digital platforms with high availability expectations. + Strong Linux administration and troubleshooting skills. + Expertise with Kubernetes and containerized environments. + Experience with AWS and/or Azure cloud platforms. + Strong scripting and automation skills in Python and Bash; Go is preferred. + Strong experience with Terraform, Git, CI/CD pipelines, Docker, and Kubernetes. + Experience working with APIs, microservices, and distributed architectures. + Hands-on experience with observability and monitoring tools such as Splunk, Grafana, Prometheus, OpenTelemetry, New Relic, and cloud-native monitoring platforms. + Demonstrated strength in incident management, escalation leadership, root cause analysis, and problem management. + Experience with capacity planning, performance engineering, disaster recovery, and resiliency testing. + Preferred: experience supporting mobile applications (iOS and Android) and digital customer platforms. + Preferred: experience with API gateways, CDN technologies, and edge architectures. + Preferred: knowledge of AI/ML platform operations. + Preferred: familiarity with cybersecurity best practices and DevSecOps. + Preferred: AWS Certified DevOps Engineer, Solutions Architect, Kubernetes, or related certifications. + Preferred: experience working within Agile and Product operating models., In accordance with federal law, only applicants who are legally authorized to work in the United States will be considered for this position. Must reside in the United States or have a valid U.S. address for residence. ## Description As a **Senior Site Reliability Engineer** , you will help build, scale, and operate the resilient platforms that power critical digital experiences across web, mobile, API, AI/ML, and customer-facing environments. This role is ideal for an engineer who thrives at the intersection of software, cloud infrastructure, automation, and operations, and who is passionate about improving reliability, observability, scalability, and developer experience. You will collaborate closely with software engineering, platform, architecture, product, and security teams to strengthen platform performance and availability, reduce operational toil, and advance intelligent, automated operations across mission-critical services. Responsibilities + Design, implement, and support highly available, scalable, and resilient platform services. + Define and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets. + Identify and address reliability risks and performance bottlenecks across distributed systems. + Lead root cause analysis for critical incidents and drive sustainable corrective actions. + Participate in production support and on-call rotations for business-critical applications. + Build and enhance observability solutions using tools such as Splunk, Grafana, Prometheus, Datadog, New Relic, and OpenTelemetry. + Create actionable dashboards, alerts, metrics, and reporting that improve operational visibility. + Drive continuous improvement in Mean Time to Detect (MTTD) and Mean Time to Resolution (MTTR). + Support and evolve shared platform capabilities used across multiple engineering teams. + Develop self-service platform features, reusable services, and automation frameworks that improve developer productivity. + Partner with architecture and engineering teams to define future-state platform strategies. + Design and manage cloud-native infrastructure across AWS and Azure environments. + Implement Infrastructure as Code using Terraform, CloudFormation, Helm, and Kubernetes manifests. + Automate provisioning, deployment, configuration management, and recovery processes. + Design and optimize CI/CD pipelines to enable secure, reliable, and efficient software delivery. + Improve deployment speed and quality through automation, release validation, and deployment controls. + Support GitOps operating models and deployment automation practices. + Lead operational readiness reviews, disaster recovery exercises, and resiliency initiatives. + Establish runbooks, playbooks, and automated remediation solutions. + Drive chaos engineering and resiliency testing efforts. + Ensure alignment with enterprise operational standards and security requirements. + Leverage AI-assisted operations, observability, and incident intelligence capabilities to proactively identify and mitigate risk. + Advance intelligent platform capabilities that enhance engineering efficiency and operational excellence. ## Related Videos - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [The Best Job Search Websites of 2025](https://www.wearedevelopers.com/magazine/368-the-best-job-search-websites-of-2025)