Lead Site Reliability Engineer

Draft Kings
Boston, MA, United States
15 days ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Compensation
$148,000.0 - $185,000.0
Working hours
Regular working hours

Tech stack

Amazon Web Services Distributed Systems Reliability Engineering User Environment Management Datadog Data Logging Kubernetes Information Technology Data Analytics

Job description

As a Lead Site Reliability Engineer, you’ll set the reliability standard across our Infrastructure Engineering organization. You’ll define how we measure reliability for critical services, partnering with engineering teams to build and refine Service Level Objectives that connect infrastructure performance to the experiences our platforms support. You’ll turn complex telemetry into clear, actionable insights that help teams and senior leaders make better decisions about reliability, risk, and priorities. As an individual contributor, you’ll lead through technical expertise and influence, shaping a consistent reliability practice across the organization. What you’ll do as a Lead Site Reliability Engineer Lead and mature the Service Level Objective development process across Infrastructure Engineering, establishing clear frameworks and standards for setting meaningful reliability targets. Partner with engineering teams to design and implement Service Level Objectives, beginning with the critical user journeys each service supports and translating them into measurable indicators, targets, and error budget policies. Review and refine existing reliability objectives to keep them aligned with changing customer impact, technical dependencies, and business priorities. Connect infrastructure reliability targets to the application and platform experiences they support, making dependencies and their impact on end-user experience clear and measurable. Build reporting processes and tooling that provide a clear view of reliability across critical components, translating technical signals into actionable insights for Senior Managers and Directors. Influence reliability practices across teams through technical guidance, design reviews, mentoring, and collaboration with engineering partners. Help teams distinguish meaningful service degradation from true downtime by applying thoughtful measurement strategies to complex distributed systems. What you’ll bring, The US base salary range for this full-time position is 148,000.00 USD - 185,000.00 USD, plus bonus, equity, and benefits as applicable. Our ranges are determined by role, level, and location. The compensation information displayed on each job posting reflects the range for new hire pay rates for the position across all US locations. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training. Your recruiter can share more about the specific pay range and how that was determined during the hiring process. It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment. An employer who violates this law shall be subject to criminal penalties and civil liability.

Requirements

A Bachelor’s Degree in Computer Science or a related field, or equivalent relevant education, experience, and training. At least 7 years of experience in Site Reliability Engineering, including hands-on experience defining and operationalizing Service Level Objectives, Service Level Indicators, and error budgets at scale. Deep experience with observability platforms such as Datadog, including building dashboards, monitors, and reporting from metrics and logging pipelines. Experience connecting infrastructure-level reliability objectives to application or platform-level outcomes and evaluating how technical dependencies affect end-user experience. Strong knowledge of distributed systems and the failure modes that can make reliability measurement complex, with the ability to assess what technical signals truly represent. Proven ability to influence across engineering teams, translate reliability concepts for technical and non-technical audiences, and drive alignment without direct authority. Excellent written and verbal communication skills, including experience developing and presenting reliability reporting to Senior Managers, Directors, and cross-functional stakeholders. Working knowledge of cloud and infrastructure environments such as Amazon Web Services, Kubernetes, and on-premise systems, with the technical depth to partner effectively with the teams operating them.

About the company

We’re a publicly traded (NASDAQ: DKNG) technology company headquartered in Boston. As a regulated gaming company, you may be required to obtain a gaming license issued by the appropriate state agency as a condition of employment. Don’t worry, we’ll guide you through the process if this is relevant to your role., Job Description About Johnson Controls Johnson Controls, a global leader in thermal management, mission-critical building systems, energy efficiency, and decarbonization, helps…

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:36 min

Visualizing memory limits and isolating suspicious endpoints

Dina Matveev Dina Matveev · Europe 2026 Virtual

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · World Congress 2022

1:10 min

Exposing sensitive information through partial search logs

Dennis Schulz Dennis Schulz +1 · World Congress 2026 Europe

8:02 min

Integrating service level objectives into incident management

Diana Todea · LIVE

1:08 min

Analyzing error logs and root causes using artificial intelligence

Nishil Patel Nishil Patel · World Congress 2025

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · World Congress 2021

Videos

See all

Related articles

See all