Staff Operational Support Engineer L2
Avispa Technology
United States
21 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.thejobnetwork.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$145,600.0
Working hours
Shift work
Job source
Tech stack
Application Programming Interfaces (APIs)
Artificial Intelligence
Content Delivery Networks
Cloud Computing
Continuous Integration
Noise Reduction
DevOps
Distributed Systems
Pattern Recognition
Prometheus
Software Deployment
WebRTC
+9 more
Data Logging
Delivery Pipeline
Grafana
Backend
Kubernetes
Video Streaming
Kibana
Terraform
Microservices
Job description
- Worksite: Leading audio, video, and voice technologies company (Atlanta, GA 30308 - Onsite)
- W2 Employment, Group Medical, Dental, Vision, Life, Retirement Savings Program, PSL
- 40 hours/week, 12 Month Assignment
A leading video, audio, and voice technologies company Staff Operational Support Engineer L2 to provide operational support for 24/7 live video streaming, advertising, player, and real-time delivery platforms.
Staff Operational Support Engineer L2 Responsibilities:
- Own escalated customer issues from Level 1 Support through resolution, troubleshooting complex production incidents affecting live streams, VOD playback, ad insertion, DRM, and real-time WebRTC services; operate directly in production environments to perform configuration changes, CDN adjustments, mitigations, and emergency changes when required, while providing clear and timely customer-facing communication and leading or contributing to live incident bridges with customers, internal teams, and partners.
- Work with Infrastructure as Code as the primary mechanism for safe, auditable, and repeatable production changes, using Terraform, Helm, Kubernetes manifests, GitOps workflows, CI/CD and deployment pipelines; validate and execute infrastructure and configuration changes through codified workflows and collaborate with Engineering and DevOps to improve deployment reliability and operational safety.
- Improve operational efficiency and incident response, including AI-assisted incident triage and classification, automated runbook execution, AI-based incident pattern detection, intelligent alert correlation and noise reduction, automated or improved incident communications, accelerated troubleshooting workflows, and identification of recurring or systemic issues; drive adoption of automation-first and AI-augmented operational practices.
- Support pre-event operational readiness for critical customer events through runbook checks, monitoring coverage validation, risk identification and mitigation planning, and rehearsed incident-response strategies, and respond to critical alerts within defined SLAs for stream health, player errors, and delivery infrastructure.
- Perform and contribute to root cause analyses, document findings and corrective and preventive actions, identify recurring issues and partner with Engineering and Product teams to eliminate them, improve runbooks, operational playbooks, and knowledge bases across player, advertising, live-streaming, and real-time products, support production deployments and defect resolution, provide feedback on observability, tooling gaps, and operational risks, and serve as the operational voice during post-incident reviews., * 24/7 On-Call Rotation: Includes nights, weekends, and holidays as part of a global support model, ensuring effective handoffs between shifts and regions.
Requirements
- 5+ years of relevant experience in operational, support, or similar customer-facing roles.
- Experience supporting production video streaming platforms, OTT services, and live systems.
- Troubleshooting skills across distributed systems, including APIs, microservices, and cloud infrastructure.
- Familiarity with HLS, DASH, CMAF, WebRTC, DRM, and CDN architectures.
- Experience using monitoring, alerting, and logging tools such as Grafana, Kibana/ELK, Prometheus, and Loki to diagnose live incidents.
- Ability to correlate backend streaming metrics, player telemetry, and CDN signals to diagnose live customer issues end-to-end.
- Comfort performing controlled changes in production environments.
- Working knowledge of incident management and on-call operations.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.thejobnetwork.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
EM
Eli McGarvie
over 3 years ago
IK
Igor Khokhriakov
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
about 2 months ago
LM
Luis Minvielle
Is Software Engineering Over-Saturated?
over 2 years ago
CH
Chris Heilmann
Dev Digest 121 - AI goes offline
over 2 years ago
CH
Chris Heilmann
Dev Digest 132 - Binging WADFlix?
about 2 years ago
CH
Chris Heilmann
Dev Digest 131 - AI'm not sure about OSS
about 2 years ago