Senior Site Reliability Engineer, Infrastructure

Draft Kings
Boston, MA, United States
about 1 month ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
4 years minimum
Compensation
$128,000.0 - $160,000.0
Working hours
Regular working hours

Tech stack

Amazon Web Services Computing Platforms Build Automation Cloud Computing Software Debugging Fault Tolerance Python (Programming Language) Linux System Administration Reliability Engineering Software Engineering VMware VSphere Datadog
+11 more
Data Logging Google Cloud Autoscaling Kubernetes Infrastructure Automation Frameworks Information Technology Rancher Build Tools Azure AKS Nutanix Docker

Job description

As a Senior Site Reliability Engineer, you’ll build and scale the critical Kubernetes infrastructure that powers our platforms and services. You’ll solve complex reliability challenges across public cloud and on-premise environments, designing automation-first solutions that strengthen performance and simplify operations. You’ll help shape architectural decisions, advance stability at scale, and build tools that give our teams the confidence to move quickly and deliver reliably. What you’ll do as a Senior Site Reliability Engineer Drive stability, performance, and scalability across our global compute platform spanning multiple public clouds and on-premise environments. Build self-healing, fault-tolerant infrastructure and internal tooling that automates repetitive operational work and reduces toil for Platform and Application teams. Operate and evolve our GitOps delivery model, using Rancher Fleet, Flux, and Helm to deploy core Kubernetes services and application workloads consistently and reliably. Own Kubernetes scaling and capacity strategies using technologies including Karpenter, Horizontal Pod Autoscaler (HPA), Kubernetes Event-Driven Autoscaling (KEDA), and predictive scaling based on event and calendar data. Define and monitor service-level objectives and reliability metrics for platform components using Datadog and our logging pipeline. Strengthen our engineering practices by sharing knowledge, contributing to architectural and design discussions, and participating in an on-call rotation. What you’ll bring, The US base salary range for this full-time position is 128,000.00 USD - 160,000.00 USD, plus bonus, equity, and benefits as applicable. Our ranges are determined by role, level, and location. The compensation information displayed on each job posting reflects the range for new hire pay rates for the position across all US locations. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training. Your recruiter can share more about the specific pay range and how that was determined during the hiring process. It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment. An employer who violates this law shall be subject to criminal penalties and civil liability., Build your best future with the Johnson Controls team! Who we are: Johnson Controls is global leader in smart, healthy, and sustainable buildings. Our mission is to reimagine t…

  • 12 days ago, As a Principal Site Reliability Engineer, you’ll shape the long-term strategy for the infrastructure behind one of the most demanding platforms in sports betting and gaming. You’ll…
  • 12 hours ago +

Requirements

A Bachelor’s Degree in Computer Science or a related field, or equivalent education, experience, and training. At least 4 years of experience managing distributed cloud and on-premise environments at scale, including strong hands-on experience with Amazon Web Services; experience with Google Cloud Platform, vSphere, or Nutanix is a plus. Deep expertise in Kubernetes and container orchestration, with experience designing, scaling, and troubleshooting complex workloads. Strong software development experience using languages such as Go and Python to build automation and infrastructure tooling. Working knowledge of networking and Linux-based systems, including container runtimes such as Docker and containerd, packet-level debugging, and kernel troubleshooting. Experience with Infrastructure as Code and configuration management tools to build scalable, consistent, and repeatable infrastructure.

About the company

We’re a publicly traded (NASDAQ: DKNG) technology company headquartered in Boston. As a regulated gaming company, you may be required to obtain a gaming license issued by the appropriate state agency as a condition of employment. Don’t worry, we’ll guide you through the process if this is relevant to your role.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:47 min

Comparing declarative GitOps tooling alternatives and interface priorities

Davide Imola Davide Imola · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

3:12 min

Configuring Docker images, networking protocols, and persistent storage volumes

Francesco Ciulla Francesco Ciulla · World Congress 2024

Videos

See all

Related articles

See all