> Markdown version of [/jobs/ext/2714544-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2714544-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Bolt - **Location:** Sunnyvale, CA, United States (Remote available) - **Experience:** Expert - **Salary:** $145,000.0 - $175,000.0 - **Contract:** Permanent contract - **Skills:** Proxmox, Amazon Web Services, Microsoft Azure, Bash Shell, C++ (Programming Language), Data Centers, Linux, Fault Tolerance, Networking Hardware, Python (Programming Language), Linux System Administration, OpenShift, Reliability Engineering, Prometheus, Subsystems, System Programming, Virtualization Technology, VMware VSphere, Scripting, Grafana, Containerization, Kubernetes, Hardware Infrastructure, Docker - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/site-reliability-engineer-bolt-graphics-8349460 ## About the Role * 5-7 years' experience in managing SRE related functions * Expert-level Linux systems administration across complex, production environments (this is a core requirement). * Exceptional proficiency in Bash and Python; advanced scripting and automation skills are mandatory, not optional. * Proven ability to write maintainable automation and diagnostic tooling for large-scale systems. * Deep understanding of server hardware, storage subsystems, and datacenter operations. * Hands-on experience with virtualization platforms including Proxmox (current), VMware vSphere, and/or OpenShift. * Strong experience with containerization technologies (Docker, containerd) and orchestration platforms (Kubernetes). * Experience operating workloads in AWS and/or Microsoft Azure environments. * Experience implementing observability, monitoring, and alerting using tools such as Prometheus and Grafana. Additional Qualifications: * Familiarity with systems programming languages such as C, C++, Rust, Go, and/or Julia. * Relevant certifications such as CompTIA A+, Azure Engineer, or similar are preferred. * Active government clearance or the ability to obtain one is required. ## Description Bolt Graphics is seeking a highly experienced Site Reliability Engineer (SRE) to design, build, and operate highly reliable developer and production systems. This role is mission-critical to maintaining uptime, performance, and operational excellence across compute, storage, and networking environments. Exceptional Linux expertise and advanced automation capabilities are mandatory for success in this role. What you'll do: * Design, implement, and operate highly available, fault-tolerant infrastructure and services. * Install, maintain, and upgrade server, storage, and networking hardware in office and colocation facilities. * Continuously monitor developer and production environments and proactively remediate reliability risks. * Participate in an on-call rotation and lead incident response efforts, including rapid triage, mitigation, and post-incident root cause analysis. * Respond effectively under pressure to outages and degradation events to restore service availability. * Develop, maintain, and continuously improve automation and operational tooling using Bash and Python. * Partner closely with engineering teams to support development, testing, and production workloads at scale. ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers)