Azure Platform & Reliability Engineer

Amazon.com, Inc.
St. Louis, MO, United States
7 days ago

Role details

Contract type
Temporary to permanent
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Compensation
$145,600.0 - $166,400.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Application Performance Management Microsoft Azure Bash Shell Cloud Computing DevOps Python (Programming Language) Key Management Windows PowerShell Reliability Engineering Software Tools Software Engineering
+11 more
Datadog Scripting Cloud Platform System Cloud Monitoring GitHub Copilot Kubernetes Information Technology Github Enterprise Azure AKS Terraform Databricks

Job description

  • This senior individual contributor role embeds directly with the client’s cloud platform and engineering teams to assess, improve, and sustain the reliability, observability, and operational readiness of the organization’s most critical Azure-hosted systems., * Assess critical enterprise applications against internal reliability standards, identifying gaps in monitoring, metrics, logging, alerting, and recovery.
  • Partner with application, cloud, infrastructure, and operations teams to design and implement systemic reliability improvements across production environments.
  • Use incident history and root cause analysis findings to identify recurring issues and drive corrective and preventative solutions.
  • Implement hands-on fixes and improvements directly within Azure, including Kubernetes, identity, networking, observability tooling, and infrastructure-as-code.
  • Build and maintain runbooks, dashboards, alert standards, and reusable reliability guidance to enable owning teams after each reliability mission.
  • Design and implement automation-first solutions that reduce toil, improve deployment safety, and strengthen observability across metrics, logs, and traces.
  • Champion responsible use of AI-assisted engineering tools to accelerate reliability improvements, incident analysis, documentation, and automation., Job Title: Infrastructure Security Engineer Location-Type: Remote with required travel (minimum monthly, travel expenses covered by the client), OR Local Hybrid, Charlotte, NC (onsite as needed) Start Date: ASAP Duration: Contract, 6 months, with poten…

Requirements

  • 7 years of experience in site reliability engineering, platform engineering, cloud infrastructure, or a closely related technical role
  • Bachelor’s degree in computer science, engineering, information technology, or equivalent practical experience.
  • Hands-on Azure experience with broad, end-to-end knowledge of managed services, identity, networking, monitoring, and operational practices.
  • Kubernetes experience, with AKS strongly preferred, including containerized runtimes and enterprise hosting patterns.
  • Experience with infrastructure-as-code, preferably Terraform, and scripting in at least one language such as Python, PowerShell, or Bash.
  • Demonstrated experience entering unfamiliar production systems, assessing reliability risks, and personally implementing improvements end to end.
  • Active, practical experience using and championing AI-assisted engineering tools such as GitHub Copilot and OpenAI Codex., * Experience with SLOs, SLIs, error budgets, production readiness reviews, resilience testing, or formal incident learning programs.
  • Hands-on experience with Azure Kubernetes Service, Azure Container Registry, Azure Key Vault, Azure Monitor, managed identities, or Azure Policy.
  • Experience with GitHub Enterprise workflows, including branch protection, CODEOWNERS, dependency scanning, GitOps, Helm, or Kustomize.
  • Background in software engineering or system design, with progression into cloud, platform, or SRE-focused roles.
  • Prior experience in smaller or midsize organizations, lean cloud teams, central platform teams, or consulting environments where broad ownership was required.
  • Experience with chaos or resilience testing, FinOps practices, or cloud operating model design.
  • Strong observability experience across metrics, logs, traces, alerting, dashboards, and service health monitoring using tools such as Datadog, Azure Monitor, OpenTelemetry, or Application Insights.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on mondo.gosnaphop.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

5:34 min

Managing token budgets and enterprise usage of coding agents

Chris Heilmann +2 · LIVE

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

4:21 min

Scaling operations using Azure AI Foundry tools

Maxim Salnikov Maxim Salnikov · WWC 2025

Videos

See all

Related articles

See all