Director, Platform SRO

Versant View all jobs
New York, NY, United States
10 days ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
$150,000.0 - $200,000.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Cloud Computing Cloud Engineering Disaster Recovery Electronic Publishing Fault Tolerance Reliability Engineering Runbook Data Streaming Alwayson
+9 more
Datadog Grafana Reliability of Systems HybridCloud Cloudformation ArcSight Event Correlation Terraform Splunk New Relic (SaaS)

Job description

The Director, Platform SRO is a senior, hands-on technical leader responsible for ensuring the stability, resilience, and operational readiness of mission-critical broadcast linear, live event, and digital media platforms. Operating in high-pressure, real-time environments, this consultant leads major incident response efforts, supports on-air and live-event continuity, and partners closely with engineering, broadcast operations, production, and vendor teams to minimize service disruption and audience impact. The role requires deep practical experience with media workflows, rapid troubleshooting during live events, and the ability to make sound technical decisions under tight time constraints. Beyond reactive incident response, the Director plays a strategic role in improving long-term system reliability and operational maturity. By applying SRO/SRE principles adapted for media environments, the consultant identifies systemic risks, drives root cause analysis, strengthens monitoring and observability, and improves operational processes across broadcast and digital ecosystems. This role balances immediate hands-on execution with advisory leadership, helping organizations build more resilient architectures, clearer incident processes, and greater confidence in their ability to support live, always-on media operations. Responsibilities Lead and coordinate high-severity incident response for broadcast linear channels, live events, and digital media platforms, serving as incident commander when required Rapidly triage and troubleshoot issues across media workflows, including playout, live production, contribution/distribution, and OTT delivery Establish, refine, and execute incident management processes, including escalation models, on-call coordination, communications, and severity classification Produce post-incident reviews, root cause analyses, and corrective action plans to prevent recurrence and reduce operational risk Assess system reliability, fault tolerance, and operational readiness across on-prem, hybrid, and cloud-based media architectures Identify single points of failure and recommend architectural, workflow, and operational improvements to enhance availability and resilience Define and improve monitoring, alerting, and observability strategies tailored to real-time broadcast and live event environments Support disaster recovery, failover planning, and live-event readiness reviews, including testing and validation Develop and maintain operational runbooks, standard operating procedures, and incident documentation Partner with engineering, broadcast operations, production teams, and vendors to align reliability practices with on-air and live-event requirements Mentor teams on incident response best practices, reliability engineering concepts, and continuous improvement Advise leadership on operational risk, system health, and reliability priorities for critical media platforms

Requirements

  • Experience supporting media, broadcast, streaming, digital publishing, or other 24x7 customer-

facing platforms.

  • Experience building or scaling SRE organizations and operational maturity programs. Hands-on

experience with observability platforms such as Datadog, New Relic, Splunk, Grafana, or similar tools.

  • Familiarity with Infrastructure as Code and automation frameworks including Terraform,

CloudFormation, or equivalent technologies.

  • Experience leading reliability initiatives across hybrid cloud and on-premises environments.

  • Industry certifications such as AWS Solutions Architect, Google Professional Cloud Engineer,

Azure Solutions Architect, ITIL, SRE Foundation, or equivalent.

  • Experience implementing AI-assisted operational intelligence, event correlation, or automated

incident response capabilities.

About the company

VERSANT (Nasdaq: VSNT) is an industry-changing media and entertainment business and home to trusted brands that shape culture, inform audiences, and build lasting connections. It operates across four core markets: political news and opinion, business news and personal finance, golf, and sports and genre entertainment. These markets are served through a powerful portfolio of iconic and innovative brands, including CNBC, MS NOW, USA Network, Golf Channel, Oxygen, E!, SYFY, and Versant’s sports division USA Sports, along with complementary digital assets including Fandango, Rotten Tomatoes, GolfNow and GolfPass., + New York City, NY

  • $150,000-200,000 per year Founded in 2001 by Robert DeNiro and Jane Rosenthal, Tribeca Enterprises is a diversified media and entertainment company that owns and operates the Tribeca Festival, Tribeca Studi…

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:50 min

Introduction and the value of runbooks

Hila Fish · World Congress 2023

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:32 min

Structuring automated incident workflows between runbooks and raw models

Aram Hakobyan Aram Hakobyan +1 · World Congress 2026 Europe

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · World Congress 2025

Videos

See all

Related articles

See all