Site Reliability Engineer | Hybrid

LTD Global
Berkeley, CA, United States
20 days ago
Apply on www.wayup.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Shift work
Job source

Tech stack

C (Programming Language) Java (Programming Language) Application Programming Interfaces (APIs) Build Automation C++ (Programming Language) Command-Line Interface Computer Programming Data Centers Linux Perl (Programming Language) Python (Programming Language) Network Security
+7 more
Reliability Engineering Prometheus Computer Networking Systems Firewalls (Computer Science) Kubernetes Virtual Agents Servicenow

Job description

As a Site Reliability Engineer on the Operations Technology team, you’ll be part of a round-the-clock crew keeping a national-scale HPC facility accessible, reliable, and secure. Working from advanced monitoring and data collection systems, you’ll proactively catch issues before they escalate, triage and resolve alerts across compute, storage, and network systems, and build the automation that makes the whole environment more resilient over time. You’ll also collaborate closely with cross-functional teams to coordinate maintenance, improve tooling, and ensure the infrastructure scales smoothly as demand grows, keeping the computational power behind critical scientific research running without interruption., + Monitor and triage alerts across computer, storage, network, and facility systems in real time

  • Build automation that prevents issues before they become outages
  • Develop new tools and integrations across the monitoring pipeline (APIs * alerts * action)
  • Walk the data center floor to keep power, cooling, and environmental systems humming
  • Coordinate maintenance activities across teams and keep incidents accurately tracked
  • Dig into complex, ambiguous problems and drive them to resolution

Requirements

  • Solid Linux/command-line (SSH) chops
  • Programming/scripting experience - Python, C, C++, Perl, or Java
  • A self-starter mindset - eager to pick up Kubernetes, Prometheus/VictoriaMetrics, Alertmanager, and building management/cooling systems
  • Network security fundamentals (ACLs, firewalls)
  • Strong cross-team communication and collaboration skills, + Experience building or deploying Agentic AI / autonomous automation for technical workflows
  • ServiceNow implementation experience
  • ITSM best-practice know-how

About the company

Great Organization!

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.wayup.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · World Congress 2022

56 sec

Integrating automated approval workflows into the portal

Markus Eisele Markus Eisele · World Congress 2025

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

4:04 min

Overview of Kubernetes operators and custom resource definitions

Philipp Krenn · World Congress 2022

2:27 min

Establishing a simulated technical environment for the workflow demo

Tobias Dunn-Krahn · LIVE

Videos

See all

Related articles

See all