Site Reliability Engineer - CTJ - Poly

Microsoft
Redmond, WA, United States
1 day ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Compensation
$119,800.0 - $234,700.0
Working hours
Regular working hours

Tech stack

Clean Code Principles Java (Programming Language) Microsoft Windows Amazon Web Services Microsoft Azure Bash Shell Big Data Microsoft Online Services C Sharp (Programming Language) Cloud Computing Data Governance Data Transformation
+24 more
Disaster Recovery Apache Hadoop Infrastructure as a Service (IaaS) Python (Programming Language) Windows PowerShell Reliability Engineering Microsoft SharePoint Software Engineering Azure Service Bus Scripting Google Cloud Feature Engineering Apache Spark Kubernetes Information Technology Production Code Bicep Data Management Terraform Azure Synapse Analytics Data Pipelines Docker Key Vault Programming Languages

Job description

This role is rooted in software engineering as a reliabilitylever. You willwork withteams that deliver production code, automation, and self-healing capabilities, and partner with feature engineering teams to bake in reliability, diagnosability, security, and compliance from design through operations. You will helpoperateand evolve large-scale enterpriseapplications,and multi-petabyte data platforms where availability, resilience, and uptime aremission critical. You will amplify impact by developing engineers,setting upreliability strategies, and influencing how services are built and run across organizational boundaries.

Responsibilities

Responsibilities:

  • Write secure, high-quality code that is maintainable, scalable, and performant.

  • Architect, implement, and optimize hybrid and cloud infrastructure using Infrastructure as Code (e.g.,Containers,Bicep,Terraform, AKSetc.) to improve availability, scale, security, and operational efficiency.

  • Design and implement data governance, storage, backup, and disaster recovery for a multi-petabyte Azure environment, ensuring integrity, security, and performance.

  • Build andoperatelarge-scale data pipelines and data transformations to support analytics, governance, and operational needs.

  • Evaluate emergingengineeringtools andpractices andincorporate them into the roadmap to continuously improve efficiency, reliability, and scale.

  • Deliver automation to improve service health, manageability, reliability, telemetry, and alerting, with a focus on resiliency.

  • Create andmaintainclear technical documentation and design specifications aligned with best practices.

  • Partner with engineering, project management, and operations to evolve services andoptimizeinfrastructure in support of organizational goals.

  • Participate in an on-call rotation tooperatelive services; troubleshoot and mitigate complex issues, escalate as needed, and write post-incident reviews to share learnings.

  • Identifyopportunities for automation using scripts, pipelines,policydrivenguardrails, orAIenabledtooling to reduce manual toil and increase engineering productivity.

Requirements

Master’s Degree in Computer Science, Information Technology, or related field AND 2+ years technical experience in software engineering, network engineering, or systems administration OR Bachelor’s Degree in Computer Science, Information Technology, or related field AND 4+ years technical experience in software engineering, network engineering, or systems administration OR equivalent experience. Other requirements: Security Clearance Requirements: Candidates must be able to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include, but are not limited to the following specialized security screenings:

  • The successful candidate must have an active U.S. Government Top Secret Clearance with access to Sensitive Compartmented Information (SCI) based on a Single Scope Background Investigation (SSBI) with Polygraph. Ability to meet Microsoft, customer and/or government security screening requirementsare requiredpre-offer and post-hirefor this role. Failure tomaintainor obtain theappropriate U.S.Government clearance and/or customer screening requirements may result in employment action up to and including termination.
  • Clearance Verification: This position requires successful verification of the stated security clearance to meet federal government customer requirements. You will be asked to provide clearance verification information prior to an offer of employment.
  • Citizenship & Citizenship Verification:This position requires verification of U.S. citizenship due to citizenship-based legal restrictions. Specifically, this position supports United States federal, state, and/or local United States government agency customer and is subject to certain citizenship-based restrictions where required or permitted by applicable law. To meet this legal requirement, citizenship will be verified via a valid passport, or other approved documents, or verified US government Clearance.
  • Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud background check upon hire/transfer and every two years thereafter.

Additional or preferred qualifications: Doctorate Degree in Computer Science, Information Technology, or related field AND 3+ years technical experience in software engineering, network engineering, or systems administration OR Master’s Degree in Computer Science, Information Technology, or related field AND 6+ years technical experience in software engineering, network engineering, or systems administration OR Bachelor’s Degree in Computer Science, Information Technology, or related field AND 8+ years technical experience in software engineering, network engineering, or systems administration OR equivalent experience.

  • 4+ years of experience building, deploying, andoperatingcontainerized applications and infrastructure as code (e.g., Docker, Kubernetes, Azure Container Apps/AKS/ACI, Terraform, Azure Bicep, ARMtemplates).

  • 4+ years of experience writing andmaintainingscripts for deployment, orchestration, and automation (e.g., PowerShell, Python, Bash).

  • Experience working with large datasets, data pipelines, and data transformation patterns (batch and/or streaming).

  • Experience with one or more major cloud platforms (Azure, AWS, or Google Cloud).

  • Hands-on experience with Azure services and infrastructure (e.g., ARM templates, IaaS, VMs, Key Vault, Event Hubs,Synapse,Spark/Hadoop), or equivalent services in AWS or Google Cloud.

  • Familiarity with data pipeline and transformation tooling (e.g., Spark, Hadoop) andoperatingat scale.

  • Familiarity with large-scale Microsoft enterprise services (e.g., Microsoft 365: Exchange, SharePoint, Skype, Teams).

  • Familiarity withpetabyte-scale datasets andbuildingreliable data pipelines and transformations that support mission-critical services.

  • Proficiencyin at least one programming language (e.g., C# or Java) and scripting languages such as PowerShell, Bash, and Python.

Site Reliability Engineering IC4 - The typical base pay range for this role across the U.S. is USD $119,800 - $234,700 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $160,200 - $261,000 per year.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on diversityjobs.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

2:56 min

Provisioning a secure container infrastructure with Bicep

Matthias Falkenberg +1 · WWC 2022

2:10 min

Why organizations combine big data and machine learning

Ayon Roy · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · WWC Europe 2026

Videos

See all

Related articles

See all