Site Reliability Engineer (SRE II) in Berkeley

Energy Jobline
Berkeley, CA, United States
4 days ago
Apply on www.energyjobline.com
Prepare application

Role details

Contract type
Temporary to permanent
Employment type
Full-time (> 32 hours)
Compensation
$166,400.0
Working hours
Regular working hours

Tech stack

C (Programming Language) Java (Programming Language) Bash Shell C++ (Programming Language) Command-Line Interface Data Centers Linux DevOps Perl (Programming Language) Python (Programming Language) Linux System Administration Reliability Engineering
+5 more
Prometheus Scientific Computating Kubernetes Build Tools Servicenow

Job description

We are seeking a Site Reliability Engineer (SRE II) to support Lawrence Berkeley Laboratory’s Energy Research Scientific Computing Center (NERSC). As part of a 24x7 operations team, you’ll help maintain the reliability and performance of critical high-performance computing infrastructure that supports scientific research and discovery. This role is ideal for an operations-focused engineer with strong Linux troubleshooting skills and experience in monitoring, alerting, incident response, automation, and production support.

What You’ll Do:

  • Monitor computing, storage, network, and facility systems
  • Review and respond to production alerts and incidents
  • Troubleshoot issues across Linux, applications, storage, networking, and infrastructure
  • Develop and maintain automation, monitoring, and alerting solutions
  • Build tools and integrations that support operational workflows
  • Support incident management and documentation through ServiceNow
  • Collaborate across technical teams to improve reliability and operational processes
  • Perform periodic data center walkthroughs to verify environmental, cooling, and power systems

Requirements

  • Experience supporting Linux-based production systems in a 24x7 operations, SRE, NOC, data center, or similar environment
  • Strong Linux administration and command-line experience
  • Experience troubleshooting production issues from alert through resolution
  • Experience developing tools or automation using Python, Perl, Java, C, C++, Bash, or similar
  • Experience with monitoring, alerting, and operational support workflows

  • Experience in Site Reliability Engineering (SRE), DevOps, NOC, Systems Administration, Infrastructure Operations, or Platform Engineering
  • ServiceNow experience
  • Experience with Kubernetes, Prometheus, VictoriaMetrics, Alertmanager, or similar monitoring platforms
  • Familiarity with IT Service Management (ITSM) best practices
  • Experience supporting HPC, research computing, scientific computing, or other mission-critical environments
  • Experience developing automation tools; AI-driven automation experience is a plus

zr

Benefits & conditions

  • Permanent overnight schedule of midnight to 8:00 a.m., five days per week
  • 100% onsite in Berkeley, California
  • U.S. only
  • No third-party agencies, Corp-to-Corp (C2C), or subcontracting arrangements

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.energyjobline.com
Prepare application

Good distractions

Loading talks and stories from around this role…