Manager, Software Engineering DevOps

The Options Clearing Corporation
Chicago, IL, United States
5 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours

Tech stack

DevOps Middleware Software Engineering Cloud Platform System Delivery Pipeline Mttr 3-tier Architectures

Job description

Experteer Overview In this role you lead L1/L2 production environment support across deployments, middleware, and platform infrastructure, driving incident response quality and SLA governance. You will deliver actionable metrics and ensure rapid, well-documented triage, resolution, and post-incident learning. You’ll manage a team to maintain readiness, escalate complex incidents, and implement automation to reduce toil. This is a high-impact, operations-focused leadership role at OCC with a strong emphasis on reliability and continuous improvement. Compensation / Benefits * Lead L1/L2 incident response activities from triage to closure and post-incident reporting * Oversee technical analysis of environment incidents across deployment, middleware, and platform layers * Serve as Tier 3 escalation point for complex incidents across Platform, S&I, Security, and App Dev teams * Own end-to-end incident lifecycle including RCA and permanent fixes or workarounds * Drive post-incident reviews for P1 and P2 incidents with identified root causes and action plans * Define, publish, and enforce SLA targets across severity levels and monitor real-time compliance * Publish monthly SLA reports with trend analysis and improvement actions * Lead alert tuning and MTTR initiatives to reduce toil and improve triage efficiency * Identify recurring incident patterns and drive permanent fixes; open and track Problem records * Lead automation and tooling projects to streamline support workflows * Ensure accurate incident documentation and maintain runbooks and knowledge base * Manage a team of 6-10 L1/L2 engineers including on-call coverage and career development * Oversee talent management, performance reviews, and training plans * Collaborate with cross-functional teams to align operational policies and priorities * Maintain up-to-date environment configuration documentation and procedures Tasks * Proven team leadership experience with accountability for standards across incident types and urgencies * Strong cross-functional collaboration skills across L1/L2 Platform Security and App Dev teams * Hands-on experience in production environment supports including deployment pipelines, containers, and middleware * Ability to create, tune, and maintain monitoring alerts and runbooks independently * Excellent oral and written communication skills for leadership reporting * Analytical, judgement, and consultation skills under pressure with stakeholder management * Ability to manage multiple priorities with strong organizational discipline * Minimum 5 years in environment operations or related fields with interdisciplinary exposure Key requirements * hybrid work environment (remote up to 2 days/week) * tuition reimbursement * student loan repayment assistance * technology stipend * generous PTO and parental leave * 401k employer match

Requirements

  • for P1 and P2 incidents with identified root causes and action plans * Define, publish, and enforce SLA targets across severity levels and monitor real-time compliance * Publish monthly SLA reports with trend analysis and improvement actions * Lead alert tuning and MTTR initiatives to reduce toil and improve triage efficiency * Identify recurring incident patterns and drive permanent fixes; open and track Problem records * Lead automation and tooling projects to streamline support workflows * Ensure accurate incident documentation and maintain runbooks and knowledge base * Manage a team of 6-10 L1/L2 engineers including on-call coverage and career development * Oversee talent management, performance reviews, and training plans * Collaborate with cross-functional teams to align operational policies and priorities * Maintain up-to-date environment configuration documentation and procedures Tasks * Proven team leadership experience with accountability for standards across incident types and urgencies * Strong cross-functional collaboration skills across L1/L2 Platform Security and App Dev teams * Hands-on experience in production environment supports including deployment pipelines, containers, and middleware * Ability to create, tune, and maintain monitoring alerts and runbooks independently * Excellent oral and written communication skills for leadership reporting * Analytical, judgement, and consultation skills under pressure with stakeholder management * Ability to manage multiple priorities with strong organizational discipline * Minimum 5 years in environment operations or related fields with interdisciplinary exposure Key requirements * hybrid work environment (remote up to 2 days/week) * tuition reimbursement * student loan repayment assistance * technology stipend * generous PTO and parental leave * 401k employer match

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:18 min

Implementing routing middleware for seamless multi-fragment origination

Igor Minar Igor Minar +1 · WWC 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

3:08 min

Aligning engineering processes with core business impact metrics

Chris Riley · WWC 2021

3:08 min

Shifting software delivery bottlenecks to operations and incident response

Milin Desai Milin Desai +1 · WWC Europe 2026

4:19 min

Securing API requests with frontend interceptors and backend middlewares

Bartosz Pietrucha · JS Congress

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all