> Markdown version of [/jobs/ext/3625612-site-reliability-engineer-sre-ii-in-berkeley](https://www.wearedevelopers.com/jobs/ext/3625612-site-reliability-engineer-sre-ii-in-berkeley). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer (SRE II) in Berkeley - **Company:** Energy Jobline - **Location:** Berkeley, CA, United States - **Salary:** $166,400.0 - **Contract:** Temporary to permanent - **Skills:** C (Programming Language), Java (Programming Language), Bash Shell, C++ (Programming Language), Command-Line Interface, Data Centers, Linux, DevOps, Perl (Programming Language), Python (Programming Language), Linux System Administration, Reliability Engineering, Prometheus, Scientific Computating, Kubernetes, Build Tools, Servicenow - **Published:** October 8, 2026 - **Apply:** https://www.energyjobline.com/job/site-reliability-engineer-sre-ii-berkeley-31887365 ## About the Role * Experience supporting Linux-based production systems in a 24x7 operations, SRE, NOC, data center, or similar environment * Strong Linux administration and command-line experience * Experience troubleshooting production issues from alert through resolution * Experience developing tools or automation using Python, Perl, Java, C, C++, Bash, or similar * Experience with monitoring, alerting, and operational support workflows * Experience in Site Reliability Engineering (SRE), DevOps, NOC, Systems Administration, Infrastructure Operations, or Platform Engineering * ServiceNow experience * Experience with Kubernetes, Prometheus, VictoriaMetrics, Alertmanager, or similar monitoring platforms * Familiarity with IT Service Management (ITSM) best practices * Experience supporting HPC, research computing, scientific computing, or other mission-critical environments * Experience developing automation tools; AI-driven automation experience is a plus #zr ## Description We are seeking a Site Reliability Engineer (SRE II) to support Lawrence Berkeley Laboratory's Energy Research Scientific Computing Center (NERSC). As part of a 24x7 operations team, you'll help maintain the reliability and performance of critical high-performance computing infrastructure that supports scientific research and discovery. This role is ideal for an operations-focused engineer with strong Linux troubleshooting skills and experience in monitoring, alerting, incident response, automation, and production support. What You'll Do: * Monitor computing, storage, network, and facility systems * Review and respond to production alerts and incidents * Troubleshoot issues across Linux, applications, storage, networking, and infrastructure * Develop and maintain automation, monitoring, and alerting solutions * Build tools and integrations that support operational workflows * Support incident management and documentation through ServiceNow * Collaborate across technical teams to improve reliability and operational processes * Perform periodic data center walkthroughs to verify environmental, cooling, and power systems