Staff Site Reliability Engineer

Core Scientific
Austin, TX, United States
about 2 months ago

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Applications Architecture Configuration Management DevOps Distributed Systems Github Make (Software) Python (Programming Language) Release Management Reliability Engineering Ansible Virtualization Technology
+9 more
Workflow Management Systems Datadog Reliability of Systems HybridCloud Kubernetes Infrastructure Automation Frameworks Information Technology Hardware Infrastructure Terraform

Job description

We are seeking a capable, motivated generalist who thrives in a change-controlled, compliant environment and enjoys working across hybrid cloud and on-premises systems. This role partners closely with application architecture and peer engineering teams while contributing hands-on across platform engineering, DevOps, and SRE.

This position is expected to take ownership of complex technical initiatives and see them through to completion-balancing hands-on implementation with effective delegation and cross-team coordination.

Responsibilities

  • Lead end-to-end delivery of complex technical initiatives, from problem definition and design through implementation, rollout, and operation.
  • Own the design, implementation, and reliability of systems across hybrid cloud and on-premises environments.
  • Take accountability for technical outcomes, including system reliability, scalability, and performance in regulated, change-controlled environments.
  • Drive execution by coordinating work across engineers and teams, delegating effectively while remaining hands-on where needed.
  • Partner with application architecture and peer teams to shape system design and influence technical decisions.
  • Build, deploy, and operate infrastructure and applications using automation and infrastructure as code.
  • Implement secure, immutable infrastructure using modern tooling (e.g., Terraform, Kubernetes, Helm, Ansible).
  • Improve observability, monitoring, and incident response practices.
  • Establish and promote best practices for reliability, security, and operational excellence across teams.
  • Mentor engineers and contribute to raising the technical bar across the organization.
  • Foster open, respectful, and professional communication directly within the team as well as with co-workers/ teammates and leaders across the organization.
  • Performs other duties as assigned.

Requirements

  • Bachelor’s degree in Computer Science or a related field, 7+ years of experience, or equivalent demonstrated impact in SRE, DevOps, or Infrastructure Engineering.
  • Broad technical experience across infrastructure and distributed systems, with the ability to design effective solutions, apply appropriate patterns, and anticipate scaling, reliability, and operational challenges.
  • Strong understanding of distributed systems behavior, including application runtime characteristics, service-to-service communication, networking, and failure modes in production environments.
  • Experience operating in regulated, compliant, or change-controlled environments.
  • Experience working in hybrid environments (AWS preferred; on-premises infrastructure required).
  • Strong experience with Infrastructure as Code, configuration management, and orchestration tools (Terraform, Helm, Kustomize, Ansible).
  • Experience with Kubernetes and virtualization technologies.
  • Experience with observability platforms (e.g., Datadog), including building monitoring and alerting integrations.
  • Experience with build and release systems (e.g., GitHub Actions, Makefiles, Python tooling)., While performing the duties of this job, the employee is frequently required to sit, stand, walk, use hands, and lift up to 25 pounds.

About the company

Core Scientific is a leading provider of infrastructure for high-performance compute in North America. Our mission is to accelerate digital innovation by scaling high-value compute rapidly, efficiently, and responsibly. We transform energy into high-value compute with unmatched efficiency at scale. Core Scientific is a publicly traded company (NASDAQ: CORZ).

We power AI, HPC, and other next-generation data center workloads demanding exceptional computing power, in addition to our digital asset mining operations. Our footprint consists of 11 data center campuses across seven states, housing advanced infrastructure for our customers.

What sets us apart? We have an entrepreneurial culture, a “can-do” and collaborative attitude, and we own and control our infrastructure. These strategic advantages enable us to maintain operational excellence, increase efficiency, and rapidly deploy cutting-edge innovations developed by our team of experts.

Join us and accelerate your career alongside our groundbreaking journey. We seek smart, creative, and collaborative professionals who thrive in a fast-paced, result-driven environment. Ready to be part of something exceptional? Apply today and make an impact at Core Scientific.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · WWC 2023

5:34 min

Managing token budgets and enterprise usage of coding agents

Chris Heilmann +2 · LIVE

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · WWC 2023

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · WWC 2021

Videos

See all

Related articles

See all