Director, Platform SRO

Versant Media
New York, NY, United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
$180,000.0 - $210,000.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Cloud Computing Cloud Engineering Disaster Recovery Electronic Publishing Fault Tolerance Reliability Engineering Runbook Data Streaming Alwayson
+8 more
Datadog Grafana Reliability of Systems HybridCloud Cloudformation Terraform Splunk New Relic (SaaS)

Job description

The Director, Platform SRO is a senior, hands-on technical leader responsible for ensuring the stability, resilience, and operational readiness of mission-critical broadcast linear, live event, and digital media platforms. Operating in high-pressure, real-time environments, this consultant leads major incident response efforts, supports on-air and live-event continuity, and partners closely with engineering, broadcast operations, production, and vendor teams to minimize service disruption and audience impact. The role requires deep practical experience with media workflows, rapid troubleshooting during live events, and the ability to make sound technical decisions under tight time constraints.

Beyond reactive incident response, the Director plays a strategic role in improving long-term system reliability and operational maturity. By applying SRO/SRE principles adapted for media environments, the consultant identifies systemic risks, drives root cause analysis, strengthens monitoring and observability, and improves operational processes across broadcast and digital ecosystems. This role balances immediate hands-on execution with advisory leadership, helping organizations build more resilient architectures, clearer incident processes, and greater confidence in their ability to support live, always-on media operations.

Responsibilities

  • Lead and coordinate high-severity incident response for broadcast linear channels, live events, and digital media platforms, serving as incident commander when required
  • Rapidly triage and troubleshoot issues across media workflows, including playout, live production, contribution/distribution, and OTT delivery
  • Establish, refine, and execute incident management processes, including escalation models, on-call coordination, communications, and severity classification
  • Produce post-incident reviews, root cause analyses, and corrective action plans to prevent recurrence and reduce operational risk
  • Assess system reliability, fault tolerance, and operational readiness across on-prem, hybrid, and cloud-based media architectures
  • Identify single points of failure and recommend architectural, workflow, and operational improvements to enhance availability and resilience
  • Define and improve monitoring, alerting, and observability strategies tailored to real-time broadcast and live event environments
  • Support disaster recovery, failover planning, and live-event readiness reviews, including testing and validation
  • Develop and maintain operational runbooks, standard operating procedures, and incident documentation
  • Partner with engineering, broadcast operations, production teams, and vendors to align reliability practices with on-air and live-event requirements
  • Mentor teams on incident response best practices, reliability engineering concepts, and continuous improvement
  • Advise leadership on operational risk, system health, and reliability priorities for critical media platforms

Requirements

Do you have experience in Crisis management?, + Experience supporting media, broadcast, streaming, digital publishing, or other 24x7 customer-

  • facing platforms. *

  • Experience building or scaling SRE organizations and operational maturity programs. Hands-on
  • experience with observability platforms such as Datadog, New Relic, Splunk, Grafana, or similar
  • tools. *

  • Familiarity with Infrastructure as Code and automation frameworks including Terraform,
  • CloudFormation, or equivalent technologies. *

  • Experience leading reliability initiatives across hybrid cloud and on-premises environments.
  • Industry certifications such as AWS Solutions Architect, Google Professional Cloud Engineer,
  • Azure Solutions Architect, ITIL, SRE Foundation, or equivalent. *

  • Experience implementing AI-assisted operational intelligence, event correlation, or automated
  • incident response capabilities.

About the company

VERSANT is a leading force in news, sports and entertainment - home to iconic and trusted brands that inspire, inform, and delight audiences. Our unique combination of content, technology and services enriches the cultural fabric, igniting passions, sparking conversations, and connecting people to what they love most.

As an independent, publicly traded company, VERSANT brings together powerhouse cable networks - including USA Network, CNBC, MS NOW (formerly MSNBC), Oxygen, E!, SYFY, and Golf Channel - with dynamic digital and direct-to-consumer brands such as Fandango, Rotten Tomatoes, GolfNow, GolfPass, and SportsEngine. Together, these businesses reflect our commitment to delivering exceptional experiences across every screen and service.

VERSANT is an industry-changing media company fueled by innovation and an entrepreneurial spirit. With a strong foundation and a forward-looking vision, VERSANT empowers creativity, embraces change, and drives connection in an ever-evolving world., As part of our selection process, external candidates may be required to attend an in-person interview with a VERSANT Media employee at one of our locations prior to a hiring decision. VERSANT Media’s policy is to provide equal employment opportunities to all applicants and employees without regard to race, color, religion, creed, gender, gender identity or expression, age, national origin or ancestry, citizenship, disability, sexual orientation, marital status, pregnancy, veteran status, membership in the uniformed services, genetic information, or any other basis protected by applicable law.

For LA County and City Residents Only: VERSANT Media will consider for employment qualified applicants with criminal histories, or arrest or conviction records, in a manner consistent with relevant legal requirements, including the City of Los Angeles’ Fair Chance Initiative for Hiring Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, where applicable.

If you are a qualified individual with a disability or a disabled veteran and require support throughout the application and/or recruitment process as a result of your disability, you have the right to request a reasonable accommodation. You can submit your request to candidateaccessibility@versantmedia.com.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

2:50 min

Introduction and the value of runbooks

Hila Fish · World Congress 2023

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

Videos

See all

Related articles

See all