Senior Site Reliability Engineer

Datavant
Washington, DC, United States
3 days ago
Apply on www.dcjobsite.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Compensation
$168,000.0 - $200,000.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Amazon S3 Audit Trail Microsoft Azure Cloud Computing Continuous Integration Data Infrastructure Data Systems Cursor (Graphical User Interface Elements) DevOps Fault Tolerance
+27 more
Github Identity and Access Management Python (Programming Language) Open Source Technology Reliability Engineering Azure Machine Learning Shell Script Data Streaming Datadog Data Logging Cloud Platform System Autoscaling Snowflake Grafana Multi-Cloud HybridCloud Event Driven Architecture Data Lakes Deployment Automation Apache Kafka Machine Learning Operations Amazon Simple Queue Service (SQS) Terraform Data Pipelines Devsecops Serverless Computing Databricks

Job description

We’re looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You’ll be at the forefront of building and operating a resilient, observable, and scalable platform that enables mission-critical data and ML workloads across our organization.

This role is ideal for someone who combines a strong SRE mindset with deep cloud infrastructure and data platform experience . You’re comfortable operating at scale in a complex, hybrid cloud environment and can architect systems that balance velocity, safety, and cost. You’ll work closely with Data & ML Engineers, Data Scientists, Analysts, and App Engineering teams to build a modern data platform that is secure, self-service, and production-grade.

What You Will Do

  • Operate and Improve Databricks and Snowflake : Own Databricks & Snowflake platforms lifecycle-including automation, workspace governance, job orchestration, and cost optimization.
  • Design for Reliability : Architect resilient, scalable, and secure infrastructure across cloud environments. Drive initiatives around failover, autoscaling, chaos testing, and capacity planning.
  • Advance Observability : Build and maintain platform-wide monitoring, alerting, and logging infrastructure using Datadog and other open tooling. Define and enforce SLOs/SLAs for critical services.
  • Drive CI/CD for Data & ML : Automate deployments of data pipelines, ML workflows, and infra components using GitHub Actions , Terraform, and related IaC tooling.
  • Enable Data Flow Across Platforms : Build patterns and tooling to support inter- and intra-cloud data movement across systems like Snowflake, S3, Delta Lake, and Kafka.
  • Champion Event-Driven Architectures : Leverage cloud-native tools like EventBridge , SNS/SQS, and Lambda to build loosely coupled, scalable data systems.
  • Collaborate Across Teams : Serve as the SRE and platform partner for teams across the organization, ensuring the platform meets the needs of analytics, data science, and product use cases.
  • Contribute to Strategy : Influence engineering-wide decisions on data platform architecture , ML enablement , and data product strategy .

Requirements

  • 6+ years in SRE, platform engineering, or DevOps roles supporting data-intensive or ML-powered applications.
  • AI-native working style: daily use of Claude Code, Cursor, Copilot, or equivalent, with views on how they make a team faster.
  • Hands-on Databricks experience , including workspace setup, cluster/job management, and integration with CI/CD and data orchestration tools. Experience with Snowflake as well.
  • Deep understanding of cloud-native infrastructure on AWS (or similar), including VPCs, IAM, event-driven patterns, and serverless compute.
  • Proven expertise with observability tools (especially Datadog) and architecting platform-wide logging and monitoring solutions.
  • Strong command of CI/CD tooling , especially GitHub Actions , infrastructure-as-code (Terraform), and deployment automation for data systems.
  • Working knowledge in shell scripting and Python.
  • Experience building and supporting highly available, fault-tolerant systems .
  • Excellent communication and collaboration skills; able to work effectively across teams.

What Helps You Stand Out

  • DevSecOps mindset : Familiarity with implementing security best practices in IaC, CI/CD, secret management, and audit logging.
  • Experience with ML infrastructure tooling such as MLflow, Feature Stores, and GPU workload orchestration.
  • Strong experience in both Databricks and Snowflake in a large scale production lakehouse with cross-warehouse interoperability, e.g. Iceberg v3, Glue, etc.
  • Background in compliance-aware architecture (e.g., HIPAA, SOC 2) or regulated industries.
  • Familiarity with multi-cloud or hybrid cloud data environments ; experience with Azure.
  • Contributions to open-source infrastructure, SRE, or observability tools.

About the company

Datavant is the data collaboration platform trusted for healthcare. Guided by our mission to make the world’s health data secure, accessible and actionable, we provide critical data solutions for organizations across the healthcare ecosystem - including providers, health plans, researchers, and life sciences companies. From fulfilling a single patient’s request for their medical records to powering the AI revolution in healthcare, Datavanters are building the future of how data is connected and used to improve health.

By joining Datavant today, you’re stepping onto a driven and highly collaborative team that is passionate about creating transformative change in healthcare., Datavant is committed to a work environment free from job discrimination. We are proud to be an Equal Employment Opportunity employer and all qualified applicants will receive consideration for employment without regard to race, color, sex, sexual orientation, gender identity, religion, national origin, disability, veteran status, or other legally protected status. To learn more about our commitment, please review our EEO Commitment Statement here (https://www.datavant.com/eeo-commitment-statement) . Know Your Rights (https://www.eeoc.gov/know-your-rights-workplace-discrimination-illegal) , explore the resources available through the EEOC for more information regarding your legal rights and protections. In addition, Datavant does not and will not discharge or in any other manner discriminate against employees or applicants because they have inquired about, discussed, or disclosed their own pay.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dcjobsite.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

1:36 min

Visualizing memory limits and isolating suspicious endpoints

Dina Matveev Dina Matveev · Europe 2026 Virtual

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

2:40 min

Using GitHub primitives for internal documentation and corporate operations

Kyle Daigle · Coffee With Developers

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

Videos

See all

Related articles

See all