> Markdown version of [/jobs/ext/2642423-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2642423-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Nbcuniversal Media, LLC - **Location:** Centennial, CO, United States (Remote available) - **Experience:** Expert - **Salary:** $104,000.0 - $140,000.0 - **Contract:** Permanent contract - **Skills:** Query Performance, Amazon S3, Build Automation, Bash Shell, Cloud Computing, Databases, Continuous Integration, Data Stores, Software Debugging, Linux, DevOps, Disaster Recovery, Domain Name System (DNS), Fault Tolerance, Github, Python (Programming Language), PostgreSQL, Reliability Engineering, Software Engineering, TypeScript, S3 Bucket, Pulumi, Scripting, Transport Layer Security, Backend, Servicebus, Infrastructure Automation Frameworks, Information Technology, Cloudwatch, Amazon Simple Queue Service (SQS), Serverless Computing, Cisco, Docker - **Published:** August 11, 2026 - **Apply:** https://jobs.ashbyhq.com/meridianlink/8211781d-4be8-4cf4-a047-983366e680c4 ## About the Role * Bachelor's degree and 4-6 years of related experience or equivalent work experience. * 5+ years of experience in DevOps, site reliability, or platform operations, with significant responsibility for production systems. * 3+ years of hands-on experience with AWS, with an emphasis on serverless services (Lambda, SQS, EventBridge, CloudWatch, S3). * Strong database administration experience: PostgreSQL operations, backup and recovery, and query performance; comfort administering other data stores. * Proficiency in scripting languages such as TypeScript, Python, and bash for production automation and operational tooling. * Strong understanding of Linux, DNS, TLS, Docker, GitHub Actions, and infrastructure as code (SST, Pulumi, or Terraform). * Experience with production monitoring and alerting, incident response, and on-call ownership. ## Description As a Senior Site Reliability Engineer on our cloud engineering team, you'll keep our production environment healthy, secure, and running smoothly. This is an operations-focused role: you'll own the day-to-day administration of our AWS accounts and databases, backup posture across our data stores, and production monitoring and debugging for a fully serverless platform. Your work will span the operational side of the software development life cycle - from deployment to maintenance and updates - always striving for continuous improvement. You'll keep our infrastructure clean, easily deployable, and scalable, creating a stable operating environment for the whole team., * Own day-to-day administration across AWS services, accounts, and access, as well as database administration across PostgreSQL and our other data stores. * Own backup posture across databases, S3 buckets, and queues; verify restores regularly and maintain a tested disaster recovery plan. * Proactively monitor production - CloudWatch dashboards, metric alarms, log-based metrics, and Slack alerting - addressing operational issues before they impact users. * Lead production debugging and incident response: build and maintain runbooks, participate in the on-call rotation, and resolve queue and dead-letter-queue failures through retry, redrive, and recovery. * Continuously refine our infrastructure to ensure it is easily deployable and scalable: keep infrastructure as code (SST/Pulumi) accurate, retire unused infrastructure, and keep cost visible and justified. * Share your knowledge of production operations with the team, fostering a culture of learning and growth., Design, deploy, and maintain on-premises and cloud playout infrastructure for IP video distribution. Build automation, CI/CD pipelines, monitoring, and scalable fault-tolerant systems. Drive releases, troubleshoot broadcast incidents, mentor SREs, and provide 24/7 on-call support. Top Skills: AnsibleAWSAzureBashBroadcast TechnologiesCi/CdContainerizationGCPIp VideoJavaScriptKubernetesLinuxPerlPythonRubyStreamingTerraform Circle ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [How Cisco embraced a DevOps culture within its network engineering team](https://www.wearedevelopers.com/videos/99-how-cisco-embraced-a-devops-culture-within-its-network-engineering-team) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [DevOps at Netflix](https://www.wearedevelopers.com/videos/270-devops-at-netflix) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [The Best Job Search Websites of 2025](https://www.wearedevelopers.com/magazine/368-the-best-job-search-websites-of-2025) - [The Best Software Developer Blogs to Read](https://www.wearedevelopers.com/magazine/156-the-best-software-developer-blogs-to-read)