Site Reliability Engineer

Microsoft
San Francisco, CA, United States
3 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Compensation
$102,100.0 - $202,200.0
Working hours
Regular working hours
Job source

Tech stack

C (Programming Language) Java (Programming Language) JavaScript (Programming Language) Microsoft Windows Artificial Intelligence Microsoft Online Services C Sharp (Programming Language) C++ (Programming Language) Cloud Computing Cloud Engineering Distributed Systems Monitoring of Systems
+4 more
Python (Programming Language) Reliability Engineering Information Technology Data Analytics

Job description

Help build and operate the trusted platforms that power Microsoft 365’s most critical compliance, security, and governance services. As a member of our team, you will work at the intersection of large-scale cloud engineering, service reliability, and operational excellence, helping ensure that enterprise and government customers can rely on Microsoft services every day. You’ll collaborate with engineers across Microsoft to solve complex technical challenges, improve resiliency, and drive innovation through automation, modern cloud technologies, and data-driven operations.

As a Site Reliability Engineer, you will help design, operate, and continuously improve large-scale Microsoft 365 and Purview services that support millions of users worldwide. You will partner with software engineers, service owners, and reliability teams to monitor service health, automate operational processes, investigate production issues, and implement engineering solutions that improve availability, performance, security, and customer experience.

This opportunity will allow you to accelerate your cloud engineering expertise, develop deep knowledge of distributed systems and large-scale service operations, and build advanced skills in automation, observability, incident management, and AI-powered operational tooling. You will gain end-to-end ownership experience across the service lifecycle while helping shape the future of reliability engineering through automation, data-driven operations, and AI-enabled solutions.

Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.

Responsibilities You participate in onboarding, code/design reviews, and regular meetings with the engineering teams that develop and manage those products. You independently develop code or scripts that automate the performance of repetitive and easily scalable operations processes. You design, develop, and maintain telemetry pipelines and monitoring tools that detail operations metrics. You develop, test, troubleshoot, and implement changes to optimize code and improve products. You respond to incidents during regular on-call rotations.

Requirements

  • Bachelor’s Degree in Computer Science or related technical field AND 2+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience., Security Clearance Requirements: Candidates must be able to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include, but are not limited to the following specialized security screenings:
  • The successful candidate must have an active U.S. Government Top Secret Clearance with access to Sensitive Compartmented Information (SCI) based on a Single Scope Background Investigation (SSBI). Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. Failure to maintain or obtain the appropriate U.S. Government clearance and/or customer screening requirements may result in employment action up to and including termination.
  • Clearance Verification: This position requires successful verification of the stated security clearance to meet federal government customer requirements. You will be asked to provide clearance verification information prior to an offer of employment.
  • Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud background check upon hire/transfer and every two years thereafter.
  • Citizenship & Citizenship Verification: This position requires verification of U.S. citizenship due to citizenship-based legal restrictions. Specifically, this position supports United States federal, state, and/or local United States government agency customer and is subject to certain citizenship-based restrictions where required or permitted by applicable law. To meet this legal requirement, citizenship will be verified via a valid passport, or other approved documents, or verified US government Clearance

Preferred/Additional Qualifications

  • Master’s Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR Bachelor’s Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience.

Site Reliability Engineering IC3 - The typical base pay range for this role across the U.S. is USD $102,100 - $202,200 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $133,800 - $219,200 per year.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

2:03 min

Microsoft integrating native Unix coreutils into Windows environments

Chris Heilmann +2 · LIVE

2:27 min

Introduction to WebAssembly in a cloud computing context

Edo Edo · WWC 2024

1:32 min

Structuring platforms for new services and data analytics

Nevelina Aleksandrova · LIVE

1:24 min

Installing premium packages using the Store CLI

Noraa Junker Noraa Junker · Europe 2026 Virtual

2:51 min

Alibaba Cloud developer resources and cloud computing training

Cheng Zhang · LIVE

Videos

See all

Related articles

See all