> Markdown version of [/jobs/ext/2787864-site-reliability-engineer-in-san-francisco](https://www.wearedevelopers.com/jobs/ext/2787864-site-reliability-engineer-in-san-francisco). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer in San Francisco - **Company:** Energy Jobline - **Location:** San Francisco, CA, United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Akka (Toolkit), Amazon Web Services, ARM Architecture, Microsoft Azure, Databases, Software Debugging, Java Virtual Machine (JVM), PostgreSQL, OpenID, Public Key Infrastructure, Reliability Engineering, Prometheus, Erlang, Istio, Large Language Models, Kubernetes, Apache Kafka, Linkerd (Service Mesh), Terraform, Stream Processing, Webhooks, Dynatrace - **Published:** September 8, 2026 - **Apply:** https://www.energyjobline.com/job/site-reliability-engineer-san-francisco-31568340 ## About the Role * 7+ years in infrastructure or platform engineering, with production experience across AWS and Azure. GCP is a plus. * Experience operating Kubernetes controllers built on controller-runtime in production: CRDs, admission webhooks, finalizers, status conditions, and debugging a reconcile loop that isn't converging. You can read Go well enough to trace a reconciler to a root cause. * Infrastructure as code with real production work, Terraform, and Crossplane at the level of authoring Compositions and XRDs rather than only applying claims, delivered through Flux and Kustomize. * Operating managed Postgres (RDS, Cloud SQL, or Azure Flexible Server) in production, point-in-time restore, major-version upgrades, and moving a live database between instances inside a bounded outage window. * Observability at scale with the Prometheus operator, scrape and relabel configuration, cardinality and cost control, alert rules as code, and distributed tracing with OpenTelemetry. Running a long-term metrics store (Cortex, Mimir, Thanos) is a plus, not a requirement. * Securing Kubernetes clusters, service mesh (Linkerd or similar), OIDC/workload , mTLS, cert-manager for PKI including trust-anchor rotation, and secrets management via cloud KMS. * Production on-call experience, you've carried a pager and written up what happened afterward. * Skilled use of LLMs as a tool to sharpen your work, not to run on autopilot. * Strong written communication, we weigh this heavily in our process. Nice to have * Sizing JVM services in containers, heap versus container limits, direct memory, GC behavior, and reading a heap dump. Our platform computes JVM flags per service, and getting it wrong shows up as OOMKilled. * Operating event-sourced systems, projection lag, offsets, replay semantics, and what a journal replay does to a read model. Our own control plane is event-sourced, and so are our customers' workloads. * Messaging or streaming systems (Kafka, Pub/Sub, or similar) at production scale. * Teleport or a similar access plane, managed as code. * Distributed, stateful, or actor-based systems (Akka, Erlang/OTP). ## Related Videos - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [Bringing digital education to refugee and host communities in remote regions of Africa](https://www.wearedevelopers.com/videos/644-bringing-digital-education-to-refugee-and-host-communities-in-remote-regions-of-africa) - [Keeping applications secure by evolving OAuth 2.0 and OpenID Connect](https://www.wearedevelopers.com/videos/100152-keeping-applications-secure-by-evolving-oauth-2-0-and-openid-connect) - [Flex your Energy: Building a Cloud-Native Platform for Renewable Energy Communities](https://www.wearedevelopers.com/videos/1990-flex-your-energy-building-a-cloud-native-platform-for-renewable-energy-communities) - [Get started with securing your cloud-native Java microservices applications](https://www.wearedevelopers.com/videos/123-get-started-with-securing-your-cloud-native-java-microservices-applications) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 131 - AI'm not sure about OSS](https://www.wearedevelopers.com/magazine/472-dev-digest-131-ai-m-not-sure-about-oss) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 138 - Are you secure about this?](https://www.wearedevelopers.com/magazine/486-dev-digest-138-are-you-secure-about-this)