Site Reliability Engineer III
Role details
Job location
Tech stack
Job description
- Guides and assists others in the areas of building appropriate level designs and gaining consensus from peers where appropriate, supporting adoption of site reliability engineering best practices within your team
- Collaborates with other software engineers and teams to design, develop, test, and implement deployment and reliability approaches using automated continuous integration and continuous delivery pipelines
- Uses enterprise-authorized AI capabilities within the work environment to accelerate incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
- Implements infrastructure, configuration, and network as code for the applications and platforms in your remit
- Collaborates with technical experts, key stakeholders, and team members to resolve complex problems and proactively address issues using service level indicators and objectives before they impact customers
- Applies enterprise-authorized AI capabilities within the work environment to identify patterns in operational signals that indicate reliability risk or recurring toil, prioritizing reuse-first improvements tied to SLO outcomes.
- Familiar with availability, reliability, scalability, and solutions in their applications and works with partners to improve these outcomes iteratively
- Proactively recognizes road blocks and identifies improvements to solve business problems, including exploring new technologies where appropriate
Requirements
- Formal training or certification on site reliability engineering concepts and 3+ years applied experience
- Proficient in site reliability culture and principles and familiarity with how to implement site reliability within an application or platform
- Proficient in at least one programming language such as Python, Java/Spring Boot, and .Net
- Working knowledge of using enterprise-authorized AI capabilities within the work environment to support SRE workflows with strong validation habits and awareness of data sensitivity
- Ability to validate AI-assisted operational recommendations before applying changes, escalating when uncertain and following data sensitivity requirements
- Proficient knowledge of software applications and technical processes within a given technical discipline (e.g., Cloud, AI, Android, etc.)
- Experience in observability such as white and black box monitoring, service level objective alerting, and telemetry collection
Preferred qualifications, capabilities, and skills
- Experience with continuous integration and continuous delivery tooling
- Familiarity with container and container orchestration and troubleshooting common networking technologies and issues
Benefits & conditions
We offer a competitive total rewards package including base salary determined based on the role, experience, skill set and location. Those in eligible roles may receive commission-based pay and/or discretionary incentive compensation, paid in the form of cash and/or forfeitable equity, awarded in recognition of individual achievements and contributions. We also offer a range of benefits and programs to meet employee needs, based on eligibility. These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup childcare, tuition reimbursement, mental health support, financial coaching and more. Additional details about total compensation and benefits will be provided during the hiring process.