Systems Engineer - SRE Enablement

AutoZone, Inc.
Memphis, TN, United States
3 months ago
Apply on dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Part-time / full-time
Experience level
Experienced
Experience required
4 years minimum
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Systems Engineering Software Design Patterns DevOps Fault Tolerance Python (Programming Language) Reliability Engineering Site Reliability Engineering Practices Ansible Software Engineering Scripting Google Cloud
+7 more
Reliability of Systems Kubernetes Infrastructure Automation Frameworks Information Technology Terraform Dynatrace Golang

Job description

AutoZone’s Site Reliability Engineering (SRE) team is seeking a Systems Engineer with a focus on SRE Enablement. This position is responsible for promoting reliability and operational excellence throughout the engineering organization. The successful candidate will play a key role in establishing standards, developing shared tools, providing guidance to development teams, and cultivating a culture of reliability across our hybrid infrastructure, which includes a primary emphasis on the Google Cloud Platform (Google Cloud Platform), as well as on-premises servers and applications. The SRE Enablement Engineer collaborates closely with application, infrastructure, and architecture teams to integrate SRE best practices early in the software development lifecycle and ensure platforms adhere to rigorous production readiness standards.

Responsibilities

  • Define enterprise-wide reliability standards, Service-Level Objective (SLO) frameworks, and error budget policies.
  • Establish production readiness criteria that teams must meet prior to launching and conduct production readiness reviews across teams.
  • Own, document, and maintain the internal SRE handbook and reliability playbooks.
  • Build, maintain, and standardize shared observability platforms, specifically leveraging Dynatrace to be consumed by all engineering teams.
  • Provide templates for alerting, dashboards, and runbooks across both cloud and on-premises application workloads.
  • Participate in the incident management process, including post-mortem analysis, to continuously strengthen systemic reliability.
  • Run SRE training programs and reliability workshops for engineering teams.
  • Coach and mentor teams on SLO-based thinking and error budget management.
  • Embed proactive SRE practices and a continuous improvement mindset into the broader engineering culture.
  • Track and report reliability metrics across the enterprise, rather than just a single service.
  • Identify systemic reliability gaps and trends across cross-functional teams.
  • Report organizational reliability health to leadership and hold teams accountable to agreed-upon operational standards.
  • Act as an internal consultant during architecture and system design reviews.
  • Advise development teams on reliability design patterns (e.g., circuit breakers, retries, graceful degradation) suitable for a hybrid Google Cloud Platform and on-premises environment.
  • Engage early in new product development to influence system reliability from the outset.

Requirements

  • Bachelor’s degree in computer science, MIS, Information Technology, or a related field, or equivalent practical experience. 4 to 7 years of experience in Systems Engineering, DevOps, or SRE-related roles.

  • Deep understanding of Site Reliability Engineering principles, particularly regarding SLOs, SLIs, error budgets, and TOIL reduction.
  • Hands-on experience building, administering, and optimizing observability and APM pipelines, with a strong focus on Dynatrace.
  • Strong experience deploying and supporting workloads in Google Cloud Platform (Google Cloud Platform), as well as maintaining legacy on-premises servers and applications.
  • Experience with container orchestration platforms (e.g., Kubernetes).
  • Strong programming/scripting skills (e.g., Python, Golang, Java) and experience with IaC tools (e.g., Terraform, Ansible).
  • Exceptional communication and consulting skills, with the ability to influence architecture decisions and translate technical concepts to non-technical leadership.

Benefits & conditions

  • Competitive pay
  • Unrivaled company culture
  • Medical, dental and vision plans
  • Exclusive discounts and perks, including an AutoZone in-store discount
  • 401(k) with company match and Stock Purchase Plan
  • AutoZoners Living Well Program for free mental health support
  • Opportunities for career growth

Additional Benefits for Full-Time AutoZoners:

  • Paid time off
  • Life, and short- and long-term disability insurance options
  • Health Savings and Flexible Spending Accounts with wellness rewards
  • Tuition reimbursement

About the company

Since opening our first store in 1979, AutoZone has grown into a leading retailer and distributor of automotive parts and accessories across the Americas. Our customer-first mindset and commitment to Going the Extra Mile define who we are, for both our customers and AutoZoners. Working at AutoZone means being part of a team that values dedication, teamwork, and growth. Whether you’re helping customers or building your career, we provide tools and support to help you succeed and drive your future.

Benefits at AutoZone

AutoZone offers thoughtful benefits programs with one-on-one benefits guidance designed to improve AutoZoners’ physical, mental and financial well-being.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

1:42 min

Automating Skupper deployments using Ansible

Alex Soto Alex Soto · World Congress 2024

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

6:16 min

Event-driven Golang backend architecture and cloud deployment

Irina Branovic Irina Branovic · World Congress 2026 Europe

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

Videos

See all

Related articles

See all