> Markdown version of [/jobs/ext/3469678-manager-site-reliability-engineering](https://www.wearedevelopers.com/jobs/ext/3469678-manager-site-reliability-engineering). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Manager, Site Reliability Engineering - **Company:** Form Energy - **Location:** Berkeley, CA, United States - **Salary:** $170,246.0 - $223,455.0 - **Contract:** Permanent contract - **Skills:** Data Analysis, Cloud Computing, DevOps, Firmware, Reliability Engineering, Cyber-physical Systems, Information Technology - **Published:** September 9, 2026 - **Apply:** https://diversityjobs.com/main/sendform/8/8/28176/1/18262534?backUrl=%2Fcareer%2F18262534%2FManager-Site-Reliability-Engineering-California-Berkeley ## About the Role * 12+ years of experience supporting complex production, industrial, energy, infrastructure, automotive, robotics, or other cyber-physical systems, including technical leadership or people management experience. * Demonstrated experience leading production incident response, technical troubleshooting, escalation management, and root-cause investigation in deployed systems. * Experience leading SRE, DevOps, or production operations teams in highly automated environments, with a track record of scaling operational capacity through software, tooling, and product improvements. * Strong ability to diagnose and coordinate resolution of problems spanning software, firmware, controls, networking, hardware, and field operations. * Experience with operational monitoring, observability, alerting, remote diagnostics, reliability practices, and the operational processes needed to support deployed products. * Proven ability to lead cross-functional teams through high-priority technical issues, communicate risk and recovery plans clearly, and translate operational learnings into improvements in product reliability, diagnosability, and serviceability. * Bachelor's degree in engineering, computer science, or a related technical discipline, or equivalent practical experience. ## Description Form Energy is hiring a Manager, Site Reliability Engineer to lead the operational function responsible for maintaining the reliability, availability, and supportability of our deployed energy storage systems. This role will own the mechanisms used to monitor fleet health, respond to incidents, triage field issues, coordinate engineering escalation, and ensure new products and releases are operationally ready. As Form Energy moves from early deployments to a growing commercial fleet, a key objective of this role is to scale operational capacity through automation, tooling, and product design improvements. The successful candidate will work across software, firmware, controls, hardware, field service, and customer-facing teams to build the scalable processes and capabilities required to enable reliable fleet operations. Relocation assistance is available. What you'll do: * Lead Product Operations Assurance with a DevOps first principals approach to fleet monitoring, incident response, field triage, engineering escalation, and production support. * Build a highly automated operating model that enables the team to support a rapidly growing deployed fleet, managing operational work through automation and driving product design improvements that reduce sustaining engineering overhead. * Establish incident management processes, including severity definitions, escalation paths, incident command, communications, and post-incident reviews. * Coordinate cross-functional engineering response to complex issues spanning software, firmware, controls, networking, hardware, and site infrastructure. * Develop diagnostic playbooks, troubleshooting procedures, and operational tooling that improve first-line response and reduce dependence on individual experts. * Partner with Data, Analytics, and Cloud Applications teams to define the telemetry, dashboards, alerts, and workflows required to effectively monitor and support the fleet. * Define operational readiness requirements for new product releases and deployments, including monitoring, diagnostics, recovery procedures, escalation paths, and support documentation. * Analyze incidents and fleet data to identify recurring failure modes and drive reliability, diagnosability, and serviceability improvements back into the product. * Establish and track operational metrics such as fleet availability, incident frequency, time to detection, time to containment, time to recovery, and recurrence. * Build the team, processes, on-call model, and automation required to scale fleet operations as the installed base grows.