Operational Support Engineer (L2)

Kani Solutions
yesterday

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Senior

Job location

Tech stack

API
Artificial Intelligence
Content Delivery Networks
DevOps
Distributed Systems
Log Analysis
Pattern Recognition
Reliability Engineering
Prometheus
WebRTC
Data Logging
Large Language Models
Grafana
Mttr
Infrastructure as Code (IaC)
Backend
Containerization
Kubernetes
Live Streaming
Video Streaming
Kibana
Terraform
Docker
Microservices

Job description

  • Live Incident Leadership: Take ownership of L1 escalations; lead incident bridges, perform real-time troubleshooting, and execute controlled emergency mitigations on production systems during live events.
  • IaC & Production Operations: Safely execute, modify, and audit infrastructure changes using Terraform, Helm, Kubernetes, GitOps, and automated CI/CD pipelines.
  • AI & Automation: Utilize AI-driven tools for automated runbook execution, intelligent alert correlation, incident triage, and pattern detection to minimize MTTR.
  • Pre-Event Readiness: Perform pre-event system checks, runbook rehearsals, and risk assessments for major, high-concurrency customer events.
  • Observability & Diagnostics: Correlate backend microservices metrics, CDN logs, and player telemetry using Grafana, Prometheus, ELK/Kibana, and Loki.
  • Post-Incident & Continuous Improvement: Drive Root Cause Analysis (RCA) processes, update operational playbooks, and collaborate with Engineering to eliminate recurring defects permanently.
  • 24/7 On-Call Operations: Participate in a global 24/7 on-call rotation with strict SLA compliance.

Requirements

Experience in the video streaming domain is mandatory.

We are seeking an experienced, hands-on Operational Support Engineer (L2) to ensure the stability, performance, and 24/7 availability of our global live video streaming, ads, player, and real-time delivery platforms.

In this role, you will take full ownership of escalated production incidents, perform live mitigations on active infrastructure, and lead technical resolution bridges during major live streaming events. You will serve as the critical bridge between Support, Engineering, DevOps, and high-value clients, leveraging Infrastructure as Code (IaC) and AI-driven automation to drive operational excellence., * Experience: 5+ years in production operations, site reliability, or L2/L3 customer-facing technical support roles.

  • Video & Delivery Domain Knowledge: Deep familiarity with video streaming workflows and protocols: HLS, DASH, CMAF, WebRTC, DRM, Server-Side Ad Insertion (SSAI), and CDN architectures.
  • Infrastructure & Containerization: Solid experience with Kubernetes, Docker, Helm, Terraform, and GitOps workflows in production environments.
  • Observability Stack: Hands-on expertise with logging and telemetry tools like Grafana, Kibana/ELK, Prometheus, and Loki.
  • Distributed Systems: Strong ability to trace issues across APIs, microservices, and multi-region cloud infrastructures.
  • Crisis Management: Calm, decisive execution under high-pressure, live-event scenarios with exceptional client-facing communication skills.

Preferred Qualifications

  • Experience leveraging AI/LLM tools for log analysis, script generation, or incident summary generation.
  • Background supporting large-scale live sporting or high-concurrency streaming events.

Apply for this position