Director Lead - Systems Engineering
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+3 more
Job description
We are seeking a Director to lead Site Reliability Engineering (SRE) and Production Operations. This senior leader is accountable for the operational health, stability, resilience, and availability of production applications and platforms supporting Allyâs Automotive and Insurance businesses. The role defines and executes the strategy for production operations and reliability engineering, driving continuous improvement and partnering across engineering, product, infrastructure, and business teams to deliver secure, stable, and scalable services.
At Ally, you get a startup feel, but experience the benefits of a company thatâs worked out the kinks and is fulfilling its purpose. Weâre always evolving and see that as a good thing. From owning our work to seeing its impact in the real world, our team is relentless in finding new ways technology can help make experiences better and help people. We are problem solvers, we value diverse thinking, we support one another, and we challenge ourselves to think bigger in the journey to deliver customer-obsessed tech solutions. To read more about what our tech team does, be sure to visit our tech blog at ally.tech
The Work Itself
Key Responsibilities * Lead SRE strategy and production operations for critical application platforms, ensuring availability, resiliency, recoverability, and performance targets are consistently achieved. * Own and evolve the operating model for production support, including incident, problem, and change risk management, as well as service restoration across the application portfolio. * Drive adoption of SRE practices, including service level indicators (SLIs), service level objectives (SLOs), error budgets, operational readiness, and automation-first engineering approaches. * Define the target-state SRE operating model and organization, including capacity planning, skill mix, and sourcing strategy (employees vs. contractors), to ensure sustainable 24x7 coverage aligned with business growth. * Establish and institutionalize best practices across SRE and application sustainment, creating consistent, scalable standards for reliability engineering and operational execution. * Establish and monitor operational health metrics, using data to identify systemic risks, improve reliability, reduce incident volume, and shorten recovery times. * Provide executive leadership during major incidents, ensuring rapid coordination, clear communication, timely escalation, and durable corrective actions. * Lead post-incident reviews and problem management efforts to resolve root causes, eliminate repeat issues, and strengthen operational discipline. * Partner with product, engineering, infrastructure, and architecture teams to embed reliability, operability, and supportability into design, delivery, and release processes. * Influence senior leaders across engineering, infrastructure, and business functions-including peer organizations and one level above-to align on reliability strategy, operating models, and investment priorities. * Lead the evolution of traditional application sustainment toward a modern SRE-led model, ensuring a balanced transition that enhances reliability without disrupting critical support responsibilities. * Drive automation and tooling investments that reduce manual effort, improve observability, streamline support processes, and increase engineering efficiency. * Evaluate and quantify the impact of AI-driven operations (AIOps) and automation accelerators, driving data-informed adoption to improve reliability, efficiency, and cost outcomes. * Define standards for monitoring, alerting, logging, capacity planning, and production readiness to strengthen proactive issue detection and service resilience. * Influence cloud and platform transformation efforts by clarifying operational ownership, improving support models, and aligning reliability practices with modern engineering patterns. * Establish a clear point of view on centralized versus distributed SRE models, shaping organizational design decisions that balance scale, accountability, and alignment with Agile delivery teams. * Build, lead, and develop high-performing teams, fostering accountability, technical depth, and a culture of continuous improvement. * Provide strategic guidance and technical assessments to senior leadership, translating operational risk and technology opportunities into clear, actionable business decisions. * Oversee vendor and partner relationships supporting production operations, ensuring service quality, accountability, and alignment with enterprise standards. * Champion operational excellence by challenging legacy practices, advancing reliability engineering maturity, and promoting modern support models. * Provide leadership accountability for 24x7 production operations, including direct engagement in major incidents and crisis events, demonstrating experience operating in high-availability, always-on environments.
What Success Looks Like
Success in this role is defined by stronger operational resilience, modernized reliability practices, and a high-performing organization that delivers consistent outcomes at scale. This leader will establish a proactive, engineering-led model for production operations that improves stability, reduces risk, and enables business growth. * Application availability, resiliency, and recovery consistently meet or exceed defined service targets for critical business services. * Incident volume, repeat issues, and time to restore service are reduced through stronger operational discipline and targeted engineering improvements. * SRE practices are embedded across the portfolio, with measurable adoption of SLOs, observability standards, automation, and production readiness. * A clearly defined and scalable SRE operating model is established, with the right balance of skills, capacity, and sourcing to support long-term needs. * Operational decisions are data-driven, supported by clear health metrics, risk indicators, and executive reporting. * Cross-functional teams demonstrate stronger accountability for reliability and improved alignment between delivery velocity and operational stability. * Manual effort and operational toil are reduced through automation, improved tooling, and streamlined processes. * Measurable gains in efficiency and effectiveness are achieved through adoption of AI-driven tooling and modern operational practices. * The organization demonstrates stronger readiness for growth, change, and platform modernization. * The function is recognized as a strategic partner that improves customer experience, protects business operations, and elevates enterprise reliability maturity.
Requirements
- 9+ Years of Relevant Experience
- Bachelorâs degree in Computer Science, Information Technology, Engineering, or equivalent
Highly Preferred Qualifications * 10+ years of experience in Site Reliability Engineering, Production Operations, Infrastructure Engineering, or related technology operations disciplines. * Strong development and engineering background. * 7+ years of experience leading managers, engineers, or large technical operations teams in complex enterprise environments. * Proven experience designing and scaling SRE or production operations organizations, including defining team structures, skill requirements, and coverage models. * Experience evaluating or implementing AIOps or AI-driven operational capabilities. * Proven experience operating in 24x7, high-availability production environments, with direct involvement in incident response and service restoration during critical events. * Proven success leading production support, incident management, problem management, and operational risk reduction for business-critical applications. * Strong knowledge of reliability engineering practices, including SLOs, operational readiness, observability, resiliency, and automation. * Experience partnering across application development, infrastructure, architecture, security, and business teams to improve service stability and delivery outcomes. * Experience leveraging metrics and operational data to drive decisions and influence senior stakeholders. * Strong communication and leadership skills, with the ability to lead through ambiguity and drive enterprise-wide improvement.
Preferred Qualifications * Advanced degree in a technical or business-related discipline. * Experience leading SRE or production operations in highly regulated, high-availability, or customer-facing environments. * Hands-on experience with cloud platforms, observability tooling, infrastructure automation, and CI/CD practices. * Experience establishing or advancing SRE practices, including SLIs, SLOs, error budgets, and toil reduction. * Demonstrated success leading large-scale transformation and operating model modernization. * Proven ability to influence senior leaders across organizational boundaries without direct authority, including experience working directly with CIO, CTO, and Development leadership. * Experience transitioning from traditional application support models to modern SRE-based approaches. * Experience managing vendors, outsourced support models, or multi-team delivery partners. * Relevant certifications in cloud, IT service management, reliability engineering, or agile practices.
Benefits & conditions
Allyâs compensation program offers market-competitive base pay and pay-for-performance incentives (bonuses) based on achieving personal and company goals. But Allyâs total compensation - or total rewards - extends beyond your paycheck and is designed to support and enrich your personal and professional life, including: * Time Away: competitive holiday and flexible paid-time-off, including time off for volunteering and voting. * Planning for the Future: plan for the near and long term with an industry-leading 401K retirement savings plan with matching and company contributions, student loan and 529 educational assistance programs, tuition reimbursement, and other financial well-being programs. * Supporting your Health & Well-being: flexible health and insurance options including dental and vision, pre-tax Health Savings Account with employer contributions and a total well-being program that helps you and your family stay on track physically, socially, emotionally, and financially. * Building a Family: adoption, surrogacy, and fertility support as well as parental and caregiver leave, back-up child and adult/elder day care program and childcare discounts. * Work-Life Integration: other benefits including LifeMattersÂŽ Employee Assistance Program, subsidized and discounted Weight WatchersÂŽ program and other employee discount programs.
About the company
Ally Financial only succeeds when its people do - and thatâs more than some clichĂŠ people put on job postings. We live this stuff! We see our people as, well, people - with interests, families, friends, dreams, and causes that are all important to them. Our focus is on the health and safety of our teammates as well as work-life balance and diversity and inclusion. From generous benefits to a variety of employee resource groups, we strive to build paths that encourage employees to stretch themselves professionally. We want to help you grow, develop, and learn new things. Youâre constantly evolving, so shouldnât your opportunities be, too?, Ally Financial is a customer-centric, leading digital financial services company with passionate customer service and innovative financial solutions. We are relentlessly focused on âDoing it Rightâ and being a trusted financial-services provider to our consumer, commercial, and corporate customers. For more information, visit www.ally.com.
Ally is an equal opportunity employer committed to diversity and inclusion in the workplace. All qualified applicants will receive consideration for employment without regard to age, race, color, sex, religion, national origin, disability, sexual orientation, gender identity or expression, pregnancy status, marital status, military or veteran status, genetic disposition or any other reason protected by law.
Where permitted by applicable law, must have received or be willing to receive the COVID-19 vaccine by date of hire to be considered, if not currently employed by Ally.
We are committed to working with and providing reasonable accommodation to applicants with physical or mental disabilities. For accommodation requests, email us at work@ally.com. Ally will not discriminate against any qualified individual who is capable of performing the essential functions of the job with or without reasonable accommodation.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on dejobs.orgGood distractions
Talks and stories from around this role â technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How to Become an AI Engineer
What is Software Engineering?
Fully Remote Software Engineer Jobs
Whatâs the Difference between a Junior, Mid, and Senior Developer?