Senior Site Reliability Engineer -AI Infrastructure Operations

NSCALE, LLC
Houston, TX, United States
25 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Compensation
$170,000.0 - $265,000.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Systems Engineering Cloud Computing Data Centers Linux Distributed Systems InfiniBand Python (Programming Language) Remote Direct Memory Access Reliability Engineering Software Engineering AI Infrastructure
+4 more
High Performance Computing Kubernetes Bare Metal Build Tools

Job description

This is a senior SRE role for someone who sets the reliability bar and then pulls the rest of the team up to it. You’ll own the hardest problems on the platform: the automation other engineers build on, the services that can’t go down, and the design decisions that determine whether either holds up at scale. You’ll still carry a pager, but the real job is making sure it fires less, for everyone, over time. What You’ll Do

  • Own reliability for critical production services end to end; set the direction, not just respond to what

breaks.

  • Grow the team, not just the systems; mentor other SREs through design review, pairing, and

incident debriefs, and hold the bar that pulls everyone up to it.

  • Set the standards the rest of the team works to: the SLO framework, the incident process, and the

on-call practices that keep it sustainable.

  • Get in early on design reviews and architecture decisions, so reliability is built in rather than bolted

on after the first outage.

  • Lead the hardest incidents and the root causes nobody else can crack; turn each one into a change

that keeps it from coming back.

  • Build the tooling and automation that removes toil for the whole team, not just your own surface

Requirements

  • 6-10 years in SRE, systems engineering, or software engineering, with real ownership of production

at scale in a data center or cloud environment.

  • Strong software engineering skills (Python, Go, or similar); you build tools other engineers adopt,

not scripts that run once and rot.

  • Deep command of Linux, networking, and distributed systems, plus the judgment to know where

the real failure modes hide.

  • Hands-on with Kubernetes and virtualized or bare-metal environments; comfortable close to the

metal, not just the cloud console.

  • Experience running AI or GPU workloads, or high-performance computing (HPC); if not, the depth to

get there fast.

  • Reliability practices you put in place that outlasted you: SLOs, observability and alerting at scale,

incident process, on-call that people can actually live with.

  • A track record as the senior voice in incidents and design reviews, trusted to make the call under

pressure.

  • A habit of raising the people around you without being asked to.

Nice to Have

  • Familiarity with high-performance networking (InfiniBand, RDMA).

Benefits & conditions

A quick note on the shape of the job. This role sits close to production, so there is an on-call rotation, and some weeks are busier than others. As a senior on the team, you help set how that rotation runs and, more to the point, how we make it lighter over time. We share the load fairly, and we treat every page as a signal worth acting on rather than just an interruption. The goal is to leave the systems quieter than you found them, so each rotation asks less of the person carrying it. If that’s the kind of ownership you’re drawn to, you’ll do well here. What We Offer

  • Competitive base plus equity, reviewed every 12 months.
  • Real ownership from the start, and a direct hand in how reliability works across the platform.
  • Flexibility that treats you as an adult; we care that the work gets done, and we trust you to shape

your day.

Salary Range $170,000 - $265,000 USD. Actual compensation varies with skill set, experience, and location, and the role may be eligible for bonus and equity.

Equal Opportunities Statement

About the company

About Nscale Nscale is the GPU cloud built for AI. We run high-performance, cost-efficient infrastructure for AI-native startups and global enterprises, from bare metal up through the platform services teams actually build on. Our culture runs on ownership, accountability, and speed. We move with urgency, we tell each other the truth, and everyone here stays close to the infrastructure that makes AI work., At Nscale, we are committed to fostering an inclusive, diverse, and equitable workplace. We believe that a variety of perspectives enrich our work environment, and we encourage applications from candidates of all backgrounds, experiences, and abilities. We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · WWC 2022

1:57 min

Routing cross-rack traffic seamlessly with NCCL

Kevin Klues Kevin Klues · WWC 2025

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · WWC Europe 2026

4:04 min

Overview of Kubernetes operators and custom resource definitions

Philipp Krenn · WWC 2022

1:06 min

Empowering site reliability engineers with integrated AI agents

Osmar Matos Osmar Matos · WWC Europe 2026

Videos

See all

Related articles

See all