> Markdown version of [/jobs/ext/2732583-site-reliability-engineer-hybrid](https://www.wearedevelopers.com/jobs/ext/2732583-site-reliability-engineer-hybrid). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer | Hybrid - **Company:** LTD Global - **Location:** Berkeley, CA, United States - **Contract:** Permanent contract - **Skills:** C (Programming Language), Java (Programming Language), Application Programming Interfaces (APIs), Build Automation, C++ (Programming Language), Command-Line Interface, Computer Programming, Data Centers, Linux, Perl (Programming Language), Python (Programming Language), Network Security, Reliability Engineering, Prometheus, Computer Networking Systems, Firewalls (Computer Science), Kubernetes, Virtual Agents, Servicenow - **Published:** September 5, 2026 - **Apply:** https://www.wayup.com/i-j-Site-Reliability-Engineer-Hybrid-LTD-GLOBAL-LLC-956562502827142/ ## About the Role + Solid Linux/command-line (SSH) chops + Programming/scripting experience - Python, C, C++, Perl, or Java + A self-starter mindset - eager to pick up Kubernetes, Prometheus/VictoriaMetrics, Alertmanager, and building management/cooling systems + Network security fundamentals (ACLs, firewalls) + Strong cross-team communication and collaboration skills, + Experience building or deploying Agentic AI / autonomous automation for technical workflows + ServiceNow implementation experience + ITSM best-practice know-how ## Description As a Site Reliability Engineer on the Operations Technology team, you'll be part of a round-the-clock crew keeping a national-scale HPC facility accessible, reliable, and secure. Working from advanced monitoring and data collection systems, you'll proactively catch issues before they escalate, triage and resolve alerts across compute, storage, and network systems, and build the automation that makes the whole environment more resilient over time. You'll also collaborate closely with cross-functional teams to coordinate maintenance, improve tooling, and ensure the infrastructure scales smoothly as demand grows, keeping the computational power behind critical scientific research running without interruption., + Monitor and triage alerts across computer, storage, network, and facility systems in real time + Build automation that prevents issues before they become outages + Develop new tools and integrations across the monitoring pipeline (APIs * alerts * action) + Walk the data center floor to keep power, cooling, and environmental systems humming + Coordinate maintenance activities across teams and keep incidents accurately tracked + Dig into complex, ambiguous problems and drive them to resolution ## Related Videos - [Applying Agile Principles to Incident Management ](https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/videos/100333-instant-kai-sandboxes-with-vcluster-multi-tenant-multi-scheduler-gpu-sharing) - [AI in Production: applied AI & enterprise use cases](https://www.wearedevelopers.com/videos/100130-ai-in-production-applied-ai-enterprise-use-cases) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)