Site Reliability Engineer

NSCALE, LLC
United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Compensation
$100,000.0 - $170,000.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Systems Engineering Computer Programming Computer Networks Data Centers Distributed Systems InfiniBand Python (Programming Language) Remote Direct Memory Access Reliability Engineering Software Engineering High Performance Computing
+4 more
Grafana Kubernetes Bare Metal Operational Systems

Job description

  • Help build and improve automation, tooling, and infrastructure that supports AI workloads
  • Support the development of operational systems and platform services
  • Assist in defining and maintaining basic SLOs/SLIs and monitoring dashboards
  • Participate in incident response, troubleshooting, and post-incident reviews
  • Investigate and help resolve performance and reliability issues across systems
  • Collaborate with Engineering, Networking, and Infrastructure teams to improve system stability
  • Contribute to improving availability, scalability, and operational efficiency
  • Learn from senior engineers and grow your expertise in reliability engineering

Requirements

Do you have experience in System troubleshooting?, * 2-5 years of experience in Site Reliability Engineering, Systems Engineering, or Software Engineering in Data Center Environment

  • 2+ years programming skills (e.g., Python, Go, or similar) with interest in automation and tooling
  • Working knowledge of Linux systems, networking concepts, and distributed systems
  • Experience troubleshooting system or application issues in production environments
  • Familiarity with monitoring or observability tools (e.g., logs, metrics, dashboards)
  • Strong willingness to learn and improve reliability and operational practices
  • Ability to work in fast-paced environments and collaborate across teams

Preferred Experience

  • Exposure to cloud platforms, Kubernetes, or virtualized/bare-metal environments
  • Experience in AI, GPU workloads, or high-performance computing (HPC)
  • Basic understanding of high-performance networking concepts (e.g., InfiniBand, RDMA)
  • Exposure to production monitoring or alerting systems at small or medium scale

Benefits & conditions

  • Highly competitive package (base + equity) with reviews every 12 months.
  • Join the fastest-growing tech startup, your chance to push boundaries, collaborate with brilliant minds, and make your mark on cutting-edge AI.
  • Expect a dynamic progression plan tailored to your ambitions. Grow by trying new things, leading, challenging the status quo, and owning your impact, always with our full support.
  • Human-First Flexibility: We treat you as humans first. Our flexible workplace trusts Nscalers to deliver, giving you the autonomy to shape your day around life’s moments.

About the company

Nscale is the GPU cloud engineered for AI-purpose-built to deliver high-performance, cost-efficient infrastructure for AI-native startups and global enterprises. We enable organizations to accelerate innovation, reduce the complexity of AI development, and achieve meaningful business outcomes through scalable, sustainable compute.

Our culture is defined by ownership, accountability, and rapid innovation. We operate with urgency and transparency, and every team member contributes to building the infrastructure powering the future of AI.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · WWC 2022

1:57 min

Routing cross-rack traffic seamlessly with NCCL

Kevin Klues Kevin Klues · WWC 2025

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

4:04 min

Overview of Kubernetes operators and custom resource definitions

Philipp Krenn · WWC 2022

1:38 min

Adopting site reliability engineering practices for machine learning

Cassie Kozyrkov · WWC 2022

Videos

See all

Related articles

See all