Sr. Manager, Site Reliability Engineering

Brinker Inc
Coppell, TX, United States
3 months ago
Apply on indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source

Tech stack

Performance Tuning Reliability Engineering Software Engineering Data Analytics Performance Monitor

Job description

We are seeking a highly skilled and motivated Sr. Manager, Site Reliability Engineering to build and lead our Site Reliability Engineering capability from the ground up. In this role, you will contribute to the vision and define operating model and foundational practices needed to improve the reliability, scalability, and performance of our technology platforms and services. You will establish core SRE disciplines such as automation, observability, incident response, service level objectives, and capacity planning while partnering closely with infrastructure, development, and operations teams to embed reliability engineering across the software lifecycle. This leader will be responsible for building a high-performing team, implementing scalable processes and tooling, and driving continuous improvement that reduces operational toil, strengthens resilience, and supports the evolving needs of the business.

This role is based in Dallas (Coppell), TX and follows a hybrid schedule (3 days in office). We are currently focused on local candidates or those open to relocating to the area at their own expense. At this time, we are unable to provide sponsorship support.

Objectives

  • Build and lead the Site Reliability Engineering capability from the ground up by establishing the team, operating model, standards, tooling, and foundational processes needed to support scalable, reliable platform infrastructure and applications.
  • Drive reliability, availability, and delivery performance by implementing automation, observability, incident response practices, and service level objectives in partnership with infrastructure, development, and operations teams.
  • Continuously improve system performance, resilience, and operational efficiency through proactive monitoring, root cause analysis, capacity planning, and data-driven optimization that reduces toil and supports evolving business needs.

Your Key Job Functions

  • Build, lead, and mentor a high-performing Site Reliability Engineering team, establishing clear priorities, accountability, and engineering standards to support a scalable and resilient operating model.
  • Define and implement the foundational SRE strategy, including service level objectives, reliability requirements, operating processes, and governance in partnership with infrastructure, development, and operations teams.
  • Design, implement, and maintain scalable and reliable infrastructure to support our applications and services.
  • Develop and maintain automation for deployment, monitoring, incident response, and operational workflows to reduce toil and improve consistency, speed, and reliability.
  • Lead or provide input into incident response and problem management practices, including root cause analysis, corrective actions, and prevention strategies to improve service availability and resilience.
  • Establish and optimize observability practices by gathering and analyzing metrics, logs, and system telemetry to support performance tuning, fault isolation, and proactive issue detection.
  • Partner with development and IT teams to embed reliability, testing, release discipline, and operational readiness into the software development lifecycle.
  • Gather and analyze metrics from operating systems, logs, as well as applications to assist in performance tuning and fa

Requirements

Do you have experience in Tooling?

About the company

What does it mean to be a BrinkerHead? It means creating moments that make everyone feel special - whether you’re supporting our restaurants, celebrating wins with your team, or sparking ideas that keep Guests coming back. We play like a team, take pride in our culture, and know that life’s too short not to work happy.

At Brinker’s Restaurant Support Center (RSC), every role fuels the success of our brands - Chili’s® Grill & Bar and Maggiano’s Little Italy® - and directly impacts Team Members and Guests. From bold ideas to everyday support, we help create a fun atmosphere, great food and drinks, and the kind of hospitality that keeps everyone coming back. Here, you’ll discover opportunities for career growth, belonging, wellbeing, and plenty of chances to work hard and have fun.

Brinker International is an equal opportunity employer. We’re proud to provide a welcoming, respectful environment where everyone can thrive.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:32 min

Structuring platforms for new services and data analytics

Nevelina Aleksandrova · LIVE

8:32 min

Benchmarking GitOps engine constraints for extensive multi-cluster environments

Artem Lajko · Europe 2026 Virtual

1:22 min

Overview of the Sentry error and performance monitoring platform

Priscila Oliveira · World Congress 2023

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · World Congress 2021

5:58 min

Analyzing web performance indicators through field and lab data

Ines Akrap Ines Akrap · LIVE

1:10 min

Introduction to Microsoft Fabric and data agents

Dr. Alexander Wachtel Dr. Alexander Wachtel +1 · World Congress 2025

Videos

See all

Related articles

See all