Operational Support Engineer (L2)
Role details
Job location
Tech stack
Job description
- Live Incident Leadership: Take ownership of L1 escalations; lead incident bridges, perform real-time troubleshooting, and execute controlled emergency mitigations on production systems during live events.
- IaC & Production Operations: Safely execute, modify, and audit infrastructure changes using Terraform, Helm, Kubernetes, GitOps, and automated CI/CD pipelines.
- AI & Automation: Utilize AI-driven tools for automated runbook execution, intelligent alert correlation, incident triage, and pattern detection to minimize MTTR.
- Pre-Event Readiness: Perform pre-event system checks, runbook rehearsals, and risk assessments for major, high-concurrency customer events.
- Observability & Diagnostics: Correlate backend microservices metrics, CDN logs, and player telemetry using Grafana, Prometheus, ELK/Kibana, and Loki.
- Post-Incident & Continuous Improvement: Drive Root Cause Analysis (RCA) processes, update operational playbooks, and collaborate with Engineering to eliminate recurring defects permanently.
- 24/7 On-Call Operations: Participate in a global 24/7 on-call rotation with strict SLA compliance.
Requirements
Experience in the video streaming domain is mandatory.
We are seeking an experienced, hands-on Operational Support Engineer (L2) to ensure the stability, performance, and 24/7 availability of our global live video streaming, ads, player, and real-time delivery platforms.
In this role, you will take full ownership of escalated production incidents, perform live mitigations on active infrastructure, and lead technical resolution bridges during major live streaming events. You will serve as the critical bridge between Support, Engineering, DevOps, and high-value clients, leveraging Infrastructure as Code (IaC) and AI-driven automation to drive operational excellence., * Experience: 5+ years in production operations, site reliability, or L2/L3 customer-facing technical support roles.
- Video & Delivery Domain Knowledge: Deep familiarity with video streaming workflows and protocols: HLS, DASH, CMAF, WebRTC, DRM, Server-Side Ad Insertion (SSAI), and CDN architectures.
- Infrastructure & Containerization: Solid experience with Kubernetes, Docker, Helm, Terraform, and GitOps workflows in production environments.
- Observability Stack: Hands-on expertise with logging and telemetry tools like Grafana, Kibana/ELK, Prometheus, and Loki.
- Distributed Systems: Strong ability to trace issues across APIs, microservices, and multi-region cloud infrastructures.
- Crisis Management: Calm, decisive execution under high-pressure, live-event scenarios with exceptional client-facing communication skills.
Preferred Qualifications
- Experience leveraging AI/LLM tools for log analysis, script generation, or incident summary generation.
- Background supporting large-scale live sporting or high-concurrency streaming events.