Director Lead - Systems Engineering .
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
Experteer Overview In this role you lead Site Reliability Engineering and Production Operations for Ally’s production platforms. You will define and execute the reliability strategy, partnering with engineering, product, and infrastructure teams to deliver secure, scalable services. You’ll drive SRE adoption, implement data-driven health metrics, and oversee 24x7 production readiness. This senior leadership role shapes the operating model, tooling, and culture to boost stability and customer experience. Compensation / Benefits * Lead SRE strategy and production operations for critical platforms to meet availability, resiliency, recoverability, and performance targets * Own and evolve the production support operating model including incident, problem, and change management and service restoration * Drive adoption of SRE practices (SLIs/SLOs, error budgets, automation-first approaches) * Define target-state SRE operating model, capacity planning, skill mix, and sourcing strategy for 24x7 coverage * Establish and standardize best practices for reliability engineering and operational execution * Monitor operational health metrics to identify risks and reduce incident volume * Lead executive communication and post-incident reviews to drive durable improvements * Collaborate with product, engineering, and architecture teams to embed reliability into design and release processes * Influence senior leaders to align reliability strategy with investment priorities * Evolve traditional application sustainment toward a modern SRE-led model with scalable practices * Invest in automation and tooling to improve observability and efficiency * Assess AI-driven operations (AIOps) and automation accelerators for value delivery * Set standards for monitoring, alerting, logging, capacity planning, and production readiness * Guide cloud and platform transformation with clear ownership and modern engineering patterns * Shape centralized vs distributed SRE models and organizational design decisions * Build and mentor high-performing, accountable teams * Provide strategic guidance to senior leadership translating risk into actionable decisions * Oversee vendor relationships to ensure service quality and enterprise alignment * Promote operational excellence and modern support models * Ensure 24x7 production operations leadership including crisis management * Support readiness for growth and platform modernization Tasks * 9+ years of relevant experience * Bachelor’s degree in Computer Science, Information Technology, Engineering, or equivalent * Strong development and engineering background * 7+ years leading managers or large technical operations teams * Experience designing and scaling SRE or production operations organizations * Experience evaluating or implementing AIOps or AI-driven operational capabilities * Experience in 24x7, high-availability production environments with incident response * Proven success in production support, incident and problem management, and risk reduction * Strong knowledge of SRE practices: SLIs/SLOs, observability, resiliency, automation * Experience partnering with cross-functional teams to improve service stability and delivery outcomes * Experience using metrics/data to influence senior stakeholders * Strong communication and leadership skills Key requirements * time off and holidays * 401K with match * health, dental, vision insurance * life insurance and disability * HSA and Healthcare FSA * parental leave and family support programs
Requirements
models * Build and mentor high-performing, accountable teams * Provide strategic guidance to senior leadership translating risk into actionable decisions * Oversee vendor relationships to ensure service quality and enterprise alignment * Promote operational excellence and modern support models * Ensure 24x7 production operations leadership including crisis management * Support readiness for growth and platform modernization Tasks * 9+ years of relevant experience * Bachelor’s degree in Computer Science, Information Technology, Engineering, or equivalent * Strong development and engineering background * 7+ years leading managers or large technical operations teams * Experience designing and scaling SRE or production operations organizations * Experience evaluating or implementing AIOps or AI-driven operational capabilities * Experience in 24x7, high-availability production environments with incident response * Proven success in production support, incident and problem management, and aa in reduction * Strong knowledge of SRE practices: SLIs/SLOs, observability, resiliency, automation * Experience partnering with cross-functional teams to improve service stability and delivery outcomes * Experience using metrics/data to influence senior stakeholders * Strong communication and leadership skills Key requirements * time off and holidays * 401K with match * health, dental, vision insurance * life insurance and disability * HSA and Healthcare FSA * parental leave and family support programs
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on us.experteer.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How to Become an AI Engineer
Fully Remote Software Engineer Jobs
How Much Does a Software Engineer Make? Realistic Software Engineering Salaries
Navigating the AI Shift