Site Reliability Engineer

Nbcuniversal Media, LLC
Centennial, CO, United States
23 days ago
Apply on jobs.ashbyhq.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Compensation
$104,000.0 - $140,000.0
Working hours
Regular working hours

Tech stack

Query Performance Amazon S3 Build Automation Bash Shell Cloud Computing Databases Continuous Integration Data Stores Software Debugging Linux DevOps Disaster Recovery
+21 more
Domain Name System (DNS) Fault Tolerance Github Python (Programming Language) PostgreSQL Reliability Engineering Software Engineering TypeScript S3 Bucket Pulumi Scripting Transport Layer Security Backend Servicebus Infrastructure Automation Frameworks Information Technology Cloudwatch Amazon Simple Queue Service (SQS) Serverless Computing Cisco Docker

Job description

As a Senior Site Reliability Engineer on our cloud engineering team, you’ll keep our production environment healthy, secure, and running smoothly. This is an operations-focused role: you’ll own the day-to-day administration of our AWS accounts and databases, backup posture across our data stores, and production monitoring and debugging for a fully serverless platform. Your work will span the operational side of the software development life cycle - from deployment to maintenance and updates - always striving for continuous improvement. You’ll keep our infrastructure clean, easily deployable, and scalable, creating a stable operating environment for the whole team., * Own day-to-day administration across AWS services, accounts, and access, as well as database administration across PostgreSQL and our other data stores.

  • Own backup posture across databases, S3 buckets, and queues; verify restores regularly and maintain a tested disaster recovery plan.
  • Proactively monitor production - CloudWatch dashboards, metric alarms, log-based metrics, and Slack alerting - addressing operational issues before they impact users.
  • Lead production debugging and incident response: build and maintain runbooks, participate in the on-call rotation, and resolve queue and dead-letter-queue failures through retry, redrive, and recovery.
  • Continuously refine our infrastructure to ensure it is easily deployable and scalable: keep infrastructure as code (SST/Pulumi) accurate, retire unused infrastructure, and keep cost visible and justified.
  • Share your knowledge of production operations with the team, fostering a culture of learning and growth., Design, deploy, and maintain on-premises and cloud playout infrastructure for IP video distribution. Build automation, CI/CD pipelines, monitoring, and scalable fault-tolerant systems. Drive releases, troubleshoot broadcast incidents, mentor SREs, and provide 24/7 on-call support. Top Skills: AnsibleAWSAzureBashBroadcast TechnologiesCi/CdContainerizationGCPIp VideoJavaScriptKubernetesLinuxPerlPythonRubyStreamingTerraform Circle

Requirements

  • Bachelor’s degree and 4-6 years of related experience or equivalent work experience.
  • 5+ years of experience in DevOps, site reliability, or platform operations, with significant responsibility for production systems.
  • 3+ years of hands-on experience with AWS, with an emphasis on serverless services (Lambda, SQS, EventBridge, CloudWatch, S3).
  • Strong database administration experience: PostgreSQL operations, backup and recovery, and query performance; comfort administering other data stores.
  • Proficiency in scripting languages such as TypeScript, Python, and bash for production automation and operational tooling.
  • Strong understanding of Linux, DNS, TLS, Docker, GitHub Actions, and infrastructure as code (SST, Pulumi, or Terraform).
  • Experience with production monitoring and alerting, incident response, and on-call ownership.

Benefits & conditions

7 Days Ago In-Office or Remote 153K-205K Annually Senior level 153K-205K Annually Senior level Blockchain * Fintech * Payments * Financial Services * Cryptocurrency * Web3 Design, build, and operate secure, scalable Kubernetes platforms and infrastructure as code (Terraform). Develop backend services and automation (Go, Python, JS/TS), improve CI/CD and observability, run on-call and incident response, define SLIs/SLOs and disaster recovery, embed security and compliance, mentor team members, and partner with product and engineering to raise reliability, performance, and cost-efficiency across hybrid and public-cloud environments. Top Skills: Ci/CdGitopsGoJavaScriptKubernetesPythonTerraformTypescript Cisco ThousandEyes, 11 Days Ago Remote or Hybrid 147K-278K Annually Senior level 147K-278K Annually Senior level Cloud * Software Design, deploy, and operate large-scale, multi-region cloud-native services to improve reliability, performance, and security. Partner with application teams to build automation, run SLO-driven incident response and on-call rotations, leverage Kubernetes and CNCF tooling, and implement scalable operations, chaos and scale testing, and infrastructure-as-code for a resilient SaaS platform. Top Skills: ArgocdAWSGoKubernetesLinux/UnixOpentelemetryPrometheusPythonService Mesh

What you need to know about the Colorado Tech Scene

With a business-friendly climate and research universities like CU Boulder and Colorado State, Colorado has made a name for itself as a startup ecosystem. The state boasts a skilled workforce and high quality of life thanks to its affordable housing, vibrant cultural scene and unparalleled opportunities for outdoor recreation. Colorado is also home to the National Renewable Energy Laboratory, helping cement its status as a hub for renewable energy innovation.

Key Facts About Colorado Tech

  • Number of Tech Workers: 260,000; 8.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Lockheed Martin, Century Link, Comcast, BAE Systems, Level 3
  • Key Industries: Software, artificial intelligence, aerospace, e-commerce, fintech, healthtech
  • Funding Landscape: $4.9 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Access Venture Partners, Ridgeline Ventures, Techstars, Blackhorn Ventures
  • Research Centers and Universities: Colorado School of Mines, University of Colorado Boulder, University of Denver, Colorado State University, Mesa Laboratory, Space Science Institute, National Center for Atmospheric Research, National Renewable Energy Laboratory, Gottlieb Institute

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on jobs.ashbyhq.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:29 min

Expanding practical knowledge with community sandboxes and resources

Stuart Clark · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

3:45 min

Prototyping deterministic agents with n8n and PyATS

Alfonso Sandoval Rosas Alfonso Sandoval Rosas · Europe 2026 Virtual

3:36 min

Critical infrastructure and performance skills for modern developers

Andrew Holway · LIVE

Videos

See all

Related articles

See all