Azure Platform & Reliability Engineer
Amazon.com, Inc.
St. Louis, MO, United States
7 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Temporary to permanent
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Compensation
$145,600.0 - $166,400.0
Working hours
Regular working hours
Job source
Tech stack
Artificial Intelligence
Application Performance Management
Microsoft Azure
Bash Shell
Cloud Computing
DevOps
Python (Programming Language)
Key Management
Windows PowerShell
Reliability Engineering
Software Tools
Software Engineering
+11 more
Datadog
Scripting
Cloud Platform System
Cloud Monitoring
GitHub Copilot
Kubernetes
Information Technology
Github Enterprise
Azure AKS
Terraform
Databricks
Job description
- This senior individual contributor role embeds directly with the client’s cloud platform and engineering teams to assess, improve, and sustain the reliability, observability, and operational readiness of the organization’s most critical Azure-hosted systems., * Assess critical enterprise applications against internal reliability standards, identifying gaps in monitoring, metrics, logging, alerting, and recovery.
- Partner with application, cloud, infrastructure, and operations teams to design and implement systemic reliability improvements across production environments.
- Use incident history and root cause analysis findings to identify recurring issues and drive corrective and preventative solutions.
- Implement hands-on fixes and improvements directly within Azure, including Kubernetes, identity, networking, observability tooling, and infrastructure-as-code.
- Build and maintain runbooks, dashboards, alert standards, and reusable reliability guidance to enable owning teams after each reliability mission.
- Design and implement automation-first solutions that reduce toil, improve deployment safety, and strengthen observability across metrics, logs, and traces.
- Champion responsible use of AI-assisted engineering tools to accelerate reliability improvements, incident analysis, documentation, and automation., Job Title: Infrastructure Security Engineer Location-Type: Remote with required travel (minimum monthly, travel expenses covered by the client), OR Local Hybrid, Charlotte, NC (onsite as needed) Start Date: ASAP Duration: Contract, 6 months, with poten…
Requirements
- 7 years of experience in site reliability engineering, platform engineering, cloud infrastructure, or a closely related technical role
- Bachelor’s degree in computer science, engineering, information technology, or equivalent practical experience.
- Hands-on Azure experience with broad, end-to-end knowledge of managed services, identity, networking, monitoring, and operational practices.
- Kubernetes experience, with AKS strongly preferred, including containerized runtimes and enterprise hosting patterns.
- Experience with infrastructure-as-code, preferably Terraform, and scripting in at least one language such as Python, PowerShell, or Bash.
- Demonstrated experience entering unfamiliar production systems, assessing reliability risks, and personally implementing improvements end to end.
- Active, practical experience using and championing AI-assisted engineering tools such as GitHub Copilot and OpenAI Codex., * Experience with SLOs, SLIs, error budgets, production readiness reviews, resilience testing, or formal incident learning programs.
- Hands-on experience with Azure Kubernetes Service, Azure Container Registry, Azure Key Vault, Azure Monitor, managed identities, or Azure Policy.
- Experience with GitHub Enterprise workflows, including branch protection, CODEOWNERS, dependency scanning, GitOps, Helm, or Kustomize.
- Background in software engineering or system design, with progression into cloud, platform, or SRE-focused roles.
- Prior experience in smaller or midsize organizations, lean cloud teams, central platform teams, or consulting environments where broad ownership was required.
- Experience with chaos or resilience testing, FinOps practices, or cloud operating model design.
- Strong observability experience across metrics, logs, traces, alerting, dashboards, and service health monitoring using tools such as Datadog, Azure Monitor, OpenTelemetry, or Application Insights.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on mondo.gosnaphop.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
AJ
Austin Joy
over 4 years ago
BR
Benjamin Ruschin
Navigating the AI Shift
11 months ago
LM
Luis Minvielle
How to Become an AI Engineer
over 2 years ago
CH
Chris Heilmann
Dev Digest 120 - Apple and peers
about 2 years ago
DC
Daniel Cranney
What is Software Engineering in the Age of AI?
10 months ago
LM
Luis Minvielle
Fully Remote Software Engineer Jobs
about 2 years ago