> Markdown version of [/jobs/ext/516776-sr-manager-site-reliability-engineering](https://www.wearedevelopers.com/jobs/ext/516776-sr-manager-site-reliability-engineering). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Sr. Manager, Site Reliability Engineering - **Company:** Brinker Inc - **Location:** Coppell, TX, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Performance Tuning, Reliability Engineering, Software Engineering, Data Analytics, Performance Monitor - **Published:** June 13, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=99250992eb5232fb ## About the Role Do you have experience in Tooling? ## Description We are seeking a highly skilled and motivated Sr. Manager, Site Reliability Engineering to build and lead our Site Reliability Engineering capability from the ground up. In this role, you will contribute to the vision and define operating model and foundational practices needed to improve the reliability, scalability, and performance of our technology platforms and services. You will establish core SRE disciplines such as automation, observability, incident response, service level objectives, and capacity planning while partnering closely with infrastructure, development, and operations teams to embed reliability engineering across the software lifecycle. This leader will be responsible for building a high-performing team, implementing scalable processes and tooling, and driving continuous improvement that reduces operational toil, strengthens resilience, and supports the evolving needs of the business. This role is based in Dallas (Coppell), TX and follows a hybrid schedule (3 days in office). We are currently focused on local candidates or those open to relocating to the area at their own expense. At this time, we are unable to provide sponsorship support. Objectives * Build and lead the Site Reliability Engineering capability from the ground up by establishing the team, operating model, standards, tooling, and foundational processes needed to support scalable, reliable platform infrastructure and applications. * Drive reliability, availability, and delivery performance by implementing automation, observability, incident response practices, and service level objectives in partnership with infrastructure, development, and operations teams. * Continuously improve system performance, resilience, and operational efficiency through proactive monitoring, root cause analysis, capacity planning, and data-driven optimization that reduces toil and supports evolving business needs. Your Key Job Functions * Build, lead, and mentor a high-performing Site Reliability Engineering team, establishing clear priorities, accountability, and engineering standards to support a scalable and resilient operating model. * Define and implement the foundational SRE strategy, including service level objectives, reliability requirements, operating processes, and governance in partnership with infrastructure, development, and operations teams. * Design, implement, and maintain scalable and reliable infrastructure to support our applications and services. * Develop and maintain automation for deployment, monitoring, incident response, and operational workflows to reduce toil and improve consistency, speed, and reliability. * Lead or provide input into incident response and problem management practices, including root cause analysis, corrective actions, and prevention strategies to improve service availability and resilience. * Establish and optimize observability practices by gathering and analyzing metrics, logs, and system telemetry to support performance tuning, fault isolation, and proactive issue detection. * Partner with development and IT teams to embed reliability, testing, release discipline, and operational readiness into the software development lifecycle. * Gather and analyze metrics from operating systems, logs, as well as applications to assist in performance tuning and fa ## Related Videos - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Practical performance tuning for Serverless Java on AWS](https://www.wearedevelopers.com/videos/2075-practical-performance-tuning-for-serverless-java-on-aws) - [Automate everything via NodeJS and Puppeteer](https://www.wearedevelopers.com/videos/322-automate-everything-via-nodejs-and-puppeteer) - [Handling incidents collaboratively is like solving a rubix cube](https://www.wearedevelopers.com/videos/680-handling-incidents-collaboratively-is-like-solving-a-rubix-cube) - [How Data is Shaping our Games](https://www.wearedevelopers.com/videos/176-how-data-is-shaping-our-games) - [Hidden efficiency - How to enable your non-engineers as a CTO](https://www.wearedevelopers.com/videos/1412-hidden-efficiency-how-to-enable-your-non-engineers-as-a-cto) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [The Best Software Developer Blogs to Read](https://www.wearedevelopers.com/magazine/156-the-best-software-developer-blogs-to-read) - [Dev Digest 139 - Soft and hard queries](https://www.wearedevelopers.com/magazine/487-dev-digest-139-soft-and-hard-queries) - [The Best X (Twitter) Accounts for Developers](https://www.wearedevelopers.com/magazine/294-the-best-x-twitter-accounts-for-developers)