> Markdown version of [/jobs/ext/2818407-senior-cloud-engineer-in-san-francisco](https://www.wearedevelopers.com/jobs/ext/2818407-senior-cloud-engineer-in-san-francisco). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Cloud Engineer in San Francisco - **Company:** Energy Jobline - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Backup Devices, Cloud Computing, Cloud Computing Security, Cloud Engineering, Continuous Integration, Software Debugging, Linux, DevOps, Disaster Recovery, Distributed Systems, Domain Name System (DNS), Firmware, Github, Identity and Access Management, Network Diagnostics, Octopus Deploy, Prometheus, Zero Trust Network Access, Software Deployment, Data Logging, Network Routers, Scripting, Cloud Platform System, Grafana, Backend, Kubernetes, Apache Kafka, Graphql, Machine Learning Operations, Terraform, Data Pipelines, Databricks - **Published:** September 10, 2026 - **Apply:** https://www.energyjobline.com/job/senior-cloud-engineer-san-francisco-31579874 ## About the Role * 5+ years of experience in DevOps, SRE, or Platform Engineering operating production AWS environments * Deep expertise with Kubernetes (EKS ), GitOps workflows (Argo CD/Flux), and Infrastructure as Code (Terraform) * Strong experience building and maintaining CI/CD pipelines, ideally with GitHub Actions * Hands-on experience operating distributed systems and cloud- platforms (e.g., Kafka/MSK) * Solid understanding of networking, DNS, TLS, /access management, and cloud security best practices * Experience with observability, monitoring, and logging tools such as Grafana, Prometheus, Loki, or similar * Strong Linux, scripting, and troubleshooting skills with the ability to debug complex production issues end-to-end Bonus Skills * Experience operating Apollo Router / GraphQL federation gateways in production. * Experience operating Argo Workflows or similar Kubernetes- job / pipeline runners in production. * Familiarity with Databricks or ML Ops pipelines for data and model deployment. * Experience designing, operating, and exercising Disaster Recovery (DR) environments, including cross-region replication, backups, and tested failover runbooks. * Experience with Tailscale or other zero-trust networking tools. * Experience supporting IoT / embedded fleets at scale, including secure device-to-cloud connectivity. * Experience in high-growth startup environments where you must wear many hats. ## Description We're scaling the deployment of critical infrastructure monitoring devices to detect real-world fault events that lead to wildfires. The platform you'll build and operate ingests millions of events per day from devices in the field, powers customer-facing dashboards and alerting, and supports the data science work that turns raw signals into grid intelligence. You will own AWS infrastructure, Kubernetes (EKS), CI/CD, and observability end-to-end, partnering with our Cloud Security team to keep the platform safe and compliant, and with backend, firmware, and data teams to keep them shipping fast. As an early member of the DevOps team, you'll have a direct hand in shaping how Gridware builds, deploys, and runs production systems for years to come. Responsibilities * Design, build, and operate scalable, secure, and highly available cloud infrastructure across AWS. * Own and evolve our Kubernetes platform, enabling reliable application deployment and operations through GitOps best practices. * Build and maintain CI/CD systems that improve developer velocity, release quality, and operational reliability. * Manage and optimize event-driven infrastructure powering high-volume telemetry and device data pipelines. * Define and maintain Infrastructure as Code standards, ensuring consistency, repeatability, and scalability across environments. * Develop and enhance observability, monitoring, and incident response capabilities to support reliable production operations. * Partner closely with Security and Engineering teams to strengthen platform security, access management, and operational resilience. * Troubleshoot complex production issues, drive root cause analysis, and turn lessons learned into automation, tooling, and operational improvements. ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Shifting Stress to Progress— Understanding DevOps to do DevOps Better](https://www.wearedevelopers.com/videos/268-shifting-stress-to-progress-understanding-devops-to-do-devops-better) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [We adopted DevOps and are Cloud-native, Now What?](https://www.wearedevelopers.com/videos/485-we-adopted-devops-and-are-cloud-native-now-what) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) ## Related Articles - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [DevOps Engineer Salary [2023]](https://www.wearedevelopers.com/magazine/203-devops-engineer-salary-2023) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)