> Markdown version of [/jobs/ext/2677103-expert-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2677103-expert-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Expert Reliability Engineer - **Company:** Cypress Semiconductor Corporation - **Location:** Reston, VA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Unit Testing, Bash Shell, Unix, Cloud Computing, Code Review, Continuous Integration, Linux, DevOps, Identity and Access Management, Python (Programming Language), Key Management, Network Security, SAP ERP, Networking Basics, Windows PowerShell, Reliability Engineering, Ansible, SAP (Applications), SAP Solution Manager, Datadog, Scripting, Cloud Platform System, Large Language Models, Grafana, Multi-Cloud, Git, Kubernetes, Terraform, Software Version Control - **Published:** September 2, 2026 - **Apply:** https://jobs.sap.com/talentcommunity/apply/1432687033/?locale=en_US ## About the Role * Ability to manage ambiguities while being innovative and collaborative * Extensive technology skills and the willingness to learn new topics quickly * Problem-solving, presentation, communication, and interpersonal skills * Ability to think strategically, delivering projects and work cross-organizationally * Knowledge of SAP and the SAP solution portfolio * Cultural awareness, intercultural competencies, and the ability to influence without formal authority * Ability to build trusted relationships with key stakeholders * Persistence, self-motivation, and willingness to work under pressure * Proven ability to work in cross-functional teams * Ability to lead and mentor junior engineers in setting and maintaining DevOps and SRE best practices * English (fluent), * 10+ years of experience in DevOps and/or SRE engineering * 7+ years of experience with a successful track record of leading engineering projects and cross-functional program teams * Deep mastery of SRE principles as they apply to a globally distributed, mission-critical platforms * Advanced Experience in architecture, engineering, and deployment of modern monitoring tooling such as Grafana, Promethius * Background in security or security-adjacent roles with a proven record of strong security foundational skills * Proven track record in end-to-end implementation initiatives * Experience in people management or staff level technical leadership * Expert level knowledge of Linux/Unix administration, networking fundamentals, and operating large-scale distributed cloud environments * Advanced proficiency with IaC, scripting, and automation such as Ansible, Terraform, Bash, Python, Powershell, CI/CD * Deep expertise with modern access control, authentication standards and identity federation at an enterprise scale. * Proven experience architecting AI-driven incident response workflows that reduce mean time to detection and resolution through automated anomaly correlation, intelligent alert triage, and LLM-assisted runbook execution * Proven ability to own the design, implementation, and continuous refinement of SRE reliability frameworks, including SLAs, SLOs, and SLIs, ensuring alignment between platform performance, business commitments, and engineering priorities ## Description In this role, you will join the Technology and Engineering team as a Site Reliability Engineer focused on securing and scaling the foundational platform that underpins SAP Sovereign Cloud. You will work alongside a globally distributed team of highly motivated engineers responsible for the design, development, deployment, and lifecycle management of the Sovereign Cloud Shared Management Services (SMS) platform, with a mandate that spans both operational reliability and platform security. As an expert Site Reliability Engineer, you will shoulder a shared ownership of the reliability, security, and operational excellence of production and non-production environments within Sovereign Cloud. It is expected that you will have extensive experience in all areas of a critical application administration stack spanning: * source control (git) * CI/CD platforms * identity and access management * secrets management * container orchestration * network security infrastructure * full-stack observability tooling across multi-cloud environments. * AI-assisted engineering workflows You will lead an experienced team of globally distributed engineers in setting standards and best practices across responsibility areas including: * Code reviews * Agentic coding (AI) security best practices * Agentic coding workflows and skill development * AI-assisted incident response * Unit test coverage * Functional test coverage * etc You will treat security as a first-class reliability concern: hardening identity and access management, secrets management, and supply chain integrity are as central to this role as uptime and incident response. You will identify and close mission-critical capability gaps, define disciplined and standardized operational processes, and help the team navigate trade-offs across deployment plans, infrastructure investments, and day-to-day operational decisions. You will own and continuously improve backup and disaster recovery drills, ensuring the platform is failure-ready at global scale. You will work closely with Sovereign Cloud operations and engineering teams, regional counterparts, and hyperscaler provider partners. ## Related Videos - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [WeAreDevelopers LIVE - Node and Package Security](https://www.wearedevelopers.com/videos/2138-wearedevelopers-live-node-and-package-security) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers) - [Top Characteristics of a Software Engineer](https://www.wearedevelopers.com/magazine/166-top-characteristics-of-a-software-engineer) - [Trustworthy AI Starts at Deployment: 5 Checks Before You Ship](https://www.wearedevelopers.com/magazine/753-trustworthy-ai-starts-at-deployment-5-checks-before-you-ship) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)