Site Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+12 more
Job description
Bolt Graphics is seeking a highly experienced Site Reliability Engineer (SRE) to design, build, and operate highly reliable developer and production systems. This role is mission-critical to maintaining uptime, performance, and operational excellence across compute, storage, and networking environments. Exceptional Linux expertise and advanced automation capabilities are mandatory for success in this role.
What you’ll do:
- Design, implement, and operate highly available, fault-tolerant infrastructure and services.
- Install, maintain, and upgrade server, storage, and networking hardware in office and colocation facilities.
- Continuously monitor developer and production environments and proactively remediate reliability risks.
- Participate in an on-call rotation and lead incident response efforts, including rapid triage, mitigation, and post-incident root cause analysis.
- Respond effectively under pressure to outages and degradation events to restore service availability.
- Develop, maintain, and continuously improve automation and operational tooling using Bash and Python.
- Partner closely with engineering teams to support development, testing, and production workloads at scale.
Requirements
- 5-7 years’ experience in managing SRE related functions
- Expert-level Linux systems administration across complex, production environments (this is a core requirement).
- Exceptional proficiency in Bash and Python; advanced scripting and automation skills are mandatory, not optional.
- Proven ability to write maintainable automation and diagnostic tooling for large-scale systems.
- Deep understanding of server hardware, storage subsystems, and datacenter operations.
- Hands-on experience with virtualization platforms including Proxmox (current), VMware vSphere, and/or OpenShift.
- Strong experience with containerization technologies (Docker, containerd) and orchestration platforms (Kubernetes).
- Experience operating workloads in AWS and/or Microsoft Azure environments.
- Experience implementing observability, monitoring, and alerting using tools such as Prometheus and Grafana.
Additional Qualifications:
- Familiarity with systems programming languages such as C, C++, Rust, Go, and/or Julia.
- Relevant certifications such as CompTIA A+, Azure Engineer, or similar are preferred.
- Active government clearance or the ability to obtain one is required.
Benefits & conditions
- Medical, Dental, & Vision - 100% covered premiums
- Equity - Stock Options
- 401(k) match
- WFH Hardware
About the company
Bolt Graphics is a semiconductor startup based in Sunnyvale, CA building the fastest and most efficient graphics processors. We pride ourselves on our first principles approach to solving problems. We are energized by our mission to reduce the barrier of entry for content creation and consumption. Our goal is to enable everyone to easily create, simulate and consume immersive experiences as vividly as they can imagine them.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Fully Remote Software Engineer Jobs
Is Software Engineering Over-Saturated?
Highest Paying Tech Companies for Developers
Why Upskilling And Reskilling is Important For Developers