> Markdown version of [/jobs/ext/2013703-director-lead-systems-engineering-organization](https://www.wearedevelopers.com/jobs/ext/2013703-director-lead-systems-engineering-organization). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Director Lead - Systems Engineering Organization - **Company:** Ally Financial Inc. - **Location:** Charlotte, NC, United States (Remote available) - **Experience:** Expert - **Salary:** $135,000.0 - $235,000.0 - **Contract:** Permanent contract - **Skills:** Agile Methodology, Artificial Intelligence, Application Portfolio Management, Systems Engineering, Cloud Computing, Cyber Security, Information Systems, Continuous Delivery, Continuous Integration, IT Management, Operational Data Store, Reliability Engineering, Site Reliability Engineering Practices, Server Administration, Software Engineering, Alwayson, Datadog, Data Logging, Cloud Platform System, System Availability, Infrastructure Automation Frameworks, Information Technology, Data Analytics - **Published:** August 10, 2026 - **Apply:** https://www.careerbuilder.com/job-details/director-lead-systems-engineering-charlotte-nc--da658acb-46ee-4d3d-8d9e-21afde1bc510 ## About the Role * 9+ Years of Relevant Experience * Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent Highly Preferred Qualifications * 10+ years of experience in Site Reliability Engineering, Production Operations, Infrastructure Engineering, or related technology operations disciplines. * Strong development and engineering background. * 7+ years of experience leading managers, engineers, or large technical operations teams in complex enterprise environments. * Proven experience designing and scaling SRE or production operations organizations, including defining team structures, skill requirements, and coverage models. * Experience evaluating or implementing AIOps or AI-driven operational capabilities. * Proven experience operating in 24x7, high-availability production environments, with direct involvement in incident response and service restoration during critical events. * Proven success leading production support, incident management, problem management, and operational risk reduction for business-critical applications. * Strong knowledge of reliability engineering practices, including SLOs, operational readiness, observability, resiliency, and automation. * Experience partnering across application development, infrastructure, architecture, security, and business teams to improve service stability and delivery outcomes. * Experience leveraging metrics and operational data to drive decisions and influence senior stakeholders. * Strong communication and leadership skills, with the ability to lead through ambiguity and drive enterprise-wide improvement., * Advanced degree in a technical or business-related discipline. * Experience leading SRE or production operations in highly regulated, high-availability, or customer-facing environments. * Hands-on experience with cloud platforms, observability tooling, infrastructure automation, and CI/CD practices. * Experience establishing or advancing SRE practices, including SLIs, SLOs, error budgets, and toil reduction. * Demonstrated success leading large-scale transformation and operating model modernization. * Proven ability to influence senior leaders across organizational boundaries without direct authority, including experience working directly with CIO, CTO, and Development leadership. * Experience transitioning from traditional application support models to modern SRE-based approaches. * Experience managing vendors, outsourced support models, or multi-team delivery partners. * Relevant certifications in cloud, IT service management, reliability engineering, or agile practices., Agile Programming Methodologies, Artificial Intelligence (AI), Auto Insurance, Automation, Best Practices, Blog, Budgeting, Business Growth, Business Operations, Business Services, Capacity Management, Channel Strategies, Channel Support, Cloud Computing, Coaching, Communication Skills, Compensation and Benefits, Computer Science, Continuous Deployment/Delivery, Continuous Improvement, Continuous Integration, Corrective Action, Cross-Functional, Customer Experience, Customer Relations, Customer Support/Service, Diversity, Embedded Systems, Exceeded Sales Goal, Financial Services, High Availability, IT Management as a Service (ITMaaS), Incident Management, Incident Response, Information Technology & Information Systems, Leadership, Machine Tool, Metrics, Operational Improvement, Operational Strategy, Operational Support, Organizational Development/Management, Outsourcing, Problem Solving Skills, Process Development, Process Improvement, Product Engineering, Production Support, Production Systems, Reliability Engineering, Risk, Risk Analysis, Risk Management, Safety/Work Safety, Security Architecture, Software Administration, Software Development, Sourcing Strategy, Standards Development, Startup, Stock Purchase Plans, Student Loans, Systems Engineering, Technical Analysis, Technical Operations, Technical Strategy, Time Management, Tuition Reimbursement, Vendor/Supplier Management ## Description * Lead SRE strategy and production operations for critical application platforms, ensuring availability, resiliency, recoverability, and performance targets are consistently achieved. * Own and evolve the operating model for production support, including incident, problem, and change risk management, as well as service restoration across the application portfolio. * Drive adoption of SRE practices, including service level indicators (SLIs), service level objectives (SLOs), error budgets, operational readiness, and automation-first engineering approaches. * Define the target-state SRE operating model and organization, including capacity planning, skill mix, and sourcing strategy (employees vs. contractors), to ensure sustainable 24x7 coverage aligned with business growth. * Establish and institutionalize best practices across SRE and application sustainment, creating consistent, scalable standards for reliability engineering and operational execution. * Establish and monitor operational health metrics, using data to identify systemic risks, improve reliability, reduce incident volume, and shorten recovery times. * Provide executive leadership during major incidents, ensuring rapid coordination, clear communication, timely escalation, and durable corrective actions. * Lead post-incident reviews and problem management efforts to resolve root causes, eliminate repeat issues, and strengthen operational discipline. * Partner with product, engineering, infrastructure, and architecture teams to embed reliability, operability, and supportability into design, delivery, and release processes. * Influence senior leaders across engineering, infrastructure, and business functions-including peer organizations and one level above-to align on reliability strategy, operating models, and investment priorities. * Lead the evolution of traditional application sustainment toward a modern SRE-led model, ensuring a balanced transition that enhances reliability without disrupting critical support responsibilities. * Drive automation and tooling investments that reduce manual effort, improve observability, streamline support processes, and increase engineering efficiency. * Evaluate and quantify the impact of AI-driven operations (AIOps) and automation accelerators, driving data-informed adoption to improve reliability, efficiency, and cost outcomes. * Define standards for monitoring, alerting, logging, capacity planning, and production readiness to strengthen proactive issue detection and service resilience. * Influence cloud and platform transformation efforts by clarifying operational ownership, improving support models, and aligning reliability practices with modern engineering patterns. * Establish a clear point of view on centralized versus distributed SRE models, shaping organizational design decisions that balance scale, accountability, and alignment with Agile delivery teams. * Build, lead, and develop high-performing teams, fostering accountability, technical depth, and a culture of continuous improvement. * Provide strategic guidance and technical assessments to senior leadership, translating operational risk and technology opportunities into clear, actionable business decisions. * Oversee vendor and partner relationships supporting production operations, ensuring service quality, accountability, and alignment with enterprise standards. * Champion operational excellence by challenging legacy practices, advancing reliability engineering maturity, and promoting modern support models. * Provide leadership accountability for 24x7 production operations, including direct engagement in major incidents and crisis events, demonstrating experience operating in high-availability, always-on environments. What Success Looks Like Success in this role is defined by stronger operational resilience, modernized reliability practices, and a high-performing organization that delivers consistent outcomes at scale. This leader will establish a proactive, engineering-led model for production operations that improves stability, reduces risk, and enables business growth. * Application availability, resiliency, and recovery consistently meet or exceed defined service targets for critical business services. * Incident volume, repeat issues, and time to restore service are reduced through stronger operational discipline and targeted engineering improvements. * SRE practices are embedded across the portfolio, with measurable adoption of SLOs, observability standards, automation, and production readiness. * A clearly defined and scalable SRE operating model is established, with the right balance of skills, capacity, and sourcing to support long-term needs. * Operational decisions are data-driven, supported by clear health metrics, risk indicators, and executive reporting. * Cross-functional teams demonstrate stronger accountability for reliability and improved alignment between delivery velocity and operational stability. * Manual effort and operational toil are reduced through automation, improved tooling, and streamlined processes. * Measurable gains in efficiency and effectiveness are achieved through adoption of AI-driven tooling and modern operational practices. * The organization demonstrates stronger readiness for growth, change, and platform modernization. * The function is recognized as a strategic partner that improves customer experience, protects business operations, and elevates enterprise reliability maturity. ## Related Videos - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) - [Crypto-secure Data Management with In-Database Blockchain](https://www.wearedevelopers.com/videos/632-crypto-secure-data-management-with-in-database-blockchain) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Build Delightful Mobile Experiences with Kotlin, Realm, and Atlas Device Sync](https://www.wearedevelopers.com/videos/694-build-delightful-mobile-experiences-with-kotlin-realm-and-atlas-device-sync) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [What is Software Engineering?](https://www.wearedevelopers.com/magazine/289-what-is-software-engineering) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers) - [Now is the time for industrialized software development](https://www.wearedevelopers.com/magazine/601-now-is-the-time-for-industrialized-software-development) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)