> Markdown version of [/jobs/ext/2709200-site-reliability-engineer-ops-automation](https://www.wearedevelopers.com/jobs/ext/2709200-site-reliability-engineer-ops-automation). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - Ops & Automation - **Company:** Cerebras Systems - **Location:** San Francisco, CA, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Continuous Delivery, Programming Tools, Python (Programming Language), Octopus Deploy, Prometheus, Delivery Pipeline, Grafana, Kubernetes, Influxdb, Build Tools, Golang - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/site-reliability-engineer-ops-automation-cerebras-ai-8852087 ## About the Role * 2-4+ years in SRE with a strong operations or automation focus * Production Kubernetes experience * Solid Python or Go for building tools and automation * Proficiency with Prometheus, Grafana, and observability-driven workflows * Ability to measure and communicate impact - reliability metrics, operational toil, velocity gains Nice-to-Have * Hands-on GitOps expertise, Argo CD / Flux or equivalent, is a plus * Experience with building continuous delivery pipelines is a strong plus * Experience with Bazel or similar build systems is a strong plus. * Familiarity with capacity planning, on-prem or multi-datacenter environments ## Description We are building a high-performance SRE function to support one of the world's fastest-growing AI inference services, powered by the Wafer-Scale Engine (WSE), helping deliver infrastructure for frontier-class models from leading model builders such as OpenAI. This role offers immediate ownership of real production systems at a growing scale, direct mentorship from seasoned engineers, and close collaboration with incoming Staff SREs who will focus on long-term automation. After ~1 month of shared hands-on operations with the Staff engineers, you'll primarily operate the current setup, bring up new capacity in high-stakes environments and help bring new continuous delivery pipelines into production use. If you thrive in high-ownership SRE roles at scale and want to help shape a team from the ground up in cutting-edge AI Inference infrastructure, this is your chance. This role does not require 24/7 on-call rotations. Key Responsibilities * Remain hands-on with operational execution (releases, capacity changes, cluster upgrades) over the next year as we build robust continuous delivery pipelines and self-service capabilities * Contribute to the development of self-service CD pipelines for key workflows using our stack: Kubernetes, Bazel, Prometheus/Grafana/InfluxDB, Python, and Go. * Build reusable automation and internal developer tools that minimize operational toil and cross-team friction * Develop and extend telemetry, observability and alerting solutions to ensure operational reliability at scale * Collaborate with Cluster Ops and development teams to identify high-impact automation opportunities and iterate quickly * Contribute to reliability practices (SLOs, post-mortems, capacity planning) ## Related Videos - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [From Factory Floor to Kubernetes Core: Building an Edge Platform One Step at a Time](https://www.wearedevelopers.com/videos/1415-from-factory-floor-to-kubernetes-core-building-an-edge-platform-one-step-at-a-time) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)