Site Reliability & Operations Engineer (L3 Support)
ActiveSoft, Inc
Blue Ash, United States of America
yesterday
Role details
Contract type
Temporary contract Employment type
Full-time (> 32 hours) Working hours
Regular working hours Languages
EnglishJob location
Blue Ash, United States of America
Tech stack
Distributed Systems
Monitoring of Systems
Reliability Engineering
Software Engineering
Grafana
Information Technology
Splunk
Appdynamics
Dynatrace
Job description
- Lead and manage Major Incident (P1/P2) response, including bridge/war-room coordination and stakeholder communication.
- Drive Root Cause Analysis (RCA), Problem Management, and corrective actions through to closure.
- Provide L3 production support for business-critical applications and infrastructure.
- Monitor, troubleshoot, and resolve production issues across cloud, on-premises, and distributed environments.
- Partner with development, infrastructure, and business teams to resolve incidents and improve platform reliability.
- Participate in application design, testing, deployment, and software delivery improvements.
- Support infrastructure and applications through a rotating 24x7 on-call schedule.
- Create and maintain operational procedures, documentation, and support processes.
Requirements
- Hands-on experience leading Major Incident Management (P1/P2) from start to resolution.
- Strong experience with Root Cause Analysis (RCA) and Problem Management methodologies (5 Whys, Fishbone, Timeline Analysis, etc.).
- Experience providing L3 Production Support or Site Reliability Engineering (SRE) support in enterprise environments.
- Strong troubleshooting skills using monitoring tools such as Dynatrace, Splunk, Grafana, AppDynamics, or similar.
- Experience supporting distributed applications across cloud and on-premises environments.
- Excellent communication and stakeholder management skills.