> Markdown version of [/jobs/ext/2704160-member-of-technical-staff-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2704160-member-of-technical-staff-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Member of Technical Staff, Site Reliability Engineer - **Company:** Inferact, Inc. - **Location:** San Francisco, CA, United States - **Salary:** $200,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Bash Shell, Cloud Computing, Computer Clusters, Continuous Integration, Software Debugging, Linux, Distributed Systems, Python (Programming Language), Reliability Engineering, Azure Machine Learning, Runbook, Service Design, Software Deployment, Scripting, Backend, Kubernetes, Information Technology, Machine Learning Operations, Terraform, Docker, Golang - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/member-of-technical-staff-site-reliability-engineer-inferact-9490140 ## About the Role * Bachelor's degree or equivalent experience in computer science, engineering, systems, infrastructure, or similar. * Strong experience operating production systems with meaningful traffic, user impact, or infrastructure criticality. * Deep understanding of SLOs, SLIs, error budgets, alerting, incident response, and post-mortem processes. * Experience live-fighting major production incidents, including mitigation, root cause analysis, escalation, and follow-through on prevention work. * Strong Linux, networking, systems debugging, observability, and distributed systems fundamentals. * Ability to design operationally simple systems and identify likely failure modes before launch. * Strong programming or scripting ability in Python, Go, Bash, or similar for automation, tooling, and reliability improvements., * Experience supporting ML infrastructure, inference systems, GPU workloads, Kubernetes-based platforms, or high-scale backend services. * Experience building or improving observability systems using metrics, logs, traces, dashboards, alerts, and runbooks. * Experience with Kubernetes, Docker, Terraform, cloud infrastructure, service meshes, CI/CD systems, or production deployment platforms. * Experience driving incident review culture, post-mortem processes, reliability reviews, and prevention-oriented engineering work. * Ability to partner with engineering teams to improve service design, release safety, capacity planning, and operational readiness. ## Description We're looking for a Site Reliability Engineer to help make vLLM-powered inference systems reliable, observable, and operationally simple at production scale. This role is for someone who thinks about failure before launch, designs systems that are easier to operate, and knows how to turn incidents into durable improvements rather than one-off fixes. You'll work across engineering and infrastructure to define SLOs, improve monitoring and alerting, strengthen incident response, drive post-mortems, and reduce operational risk before it reaches users. Your work will directly impact the reliability, availability, and production readiness of the systems powering AI inference at scale., * Owned reliability for high-throughput, latency-sensitive, or mission-critical production systems. * Supported AI inference, model serving, GPU clusters, ML platforms, or distributed serving infrastructure. * Built automation that reduced toil, improved recovery time, or prevented repeat incidents. * Led incident response for severe outages with clear communication across engineering and leadership. * Created practical SLOs, dashboards, alerts, runbooks, or release gates that improved production reliability. ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)