> Markdown version of [/jobs/ext/2720042-senior-staff-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2720042-senior-staff-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior/Staff Site Reliability Engineer - **Company:** Upstart - **Location:** Santa Clara, CA, United States - **Experience:** Expert - **Salary:** $180,000.0 - $260,000.0 - **Contract:** Permanent contract - **Skills:** Airflow, Software as a Service, Cloud Computing, Databases, Extract Transform Load (ETL), DevOps, Python (Programming Language), PostgreSQL, Package Management Systems, Reliability Engineering, Workflow Management Systems, Scripting, Grafana, Kubernetes, Low Latency, Influxdb, Deployment Automation, Docker - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/senior-staff-site-reliability-engineer-gatik-ai-7097190 ## About the Role * 5+ years of experience in a related role such as Site Reliability Engineer, DevOps Engineer, or Infrastructure Engineer. * Strong knowledge of networking fundamentals, including protocols, troubleshooting, and optimization. * Hands-on experience with Docker and related ecosystem tools (e.g., Docker Compose, Kaniko). * Expertise in Kubernetes deployments and package management via Helm. * Proficiency with relational and time-series databases (e.g., Postgres, TimescaleDB, InfluxDB). * Familiarity with workflow orchestration tools such as Argo and Airflow. * Proven experience managing upgrades and rollbacks for customer-facing SaaS environments. * Scripting experience in Python and Bash for automation and tooling. * Experience building and maintaining dashboards with tools like Grafana. ## Description We are seeking an experienced Senior/Staff Site Reliability Engineer to support the operation, monitoring, and scaling of our growing fleet of autonomous vehicles. In this role, you will work closely with our infrastructure and platform teams to manage rollouts of both on-premises and cloud infrastructure in support of expansions to new customer sites. You will be directly involved in the setup and monitoring of our data offload systems, remote supervision stations, and on-prem continuous integration (CI) environments, ensuring our infrastructure is highly reliable, secure, and optimized for performance. This position plays a critical role in keeping our autonomy operations running smoothly while supporting the rapid growth of our fleet and customer base. This role is onsite 5 days a week at our Santa Clara, CA office!, * Partner with the infrastructure and platform engineering teams to monitor, maintain, and troubleshoot our on-premises data offload and CI systems. * Design, develop, and maintain business intelligence (BI) dashboards and ETL (extract, transform, load) pipelines to provide actionable insights into our infrastructure performance and health. * Architect and deploy test environments to validate internal and customer-facing infrastructure solutions. * Automate deployment, scaling, and upgrading of our remote monitoring software to ensure operational efficiency. * Perform ongoing analysis of infrastructure performance, identifying opportunities for optimization in latency, throughput, and reliability. ## Related Videos - [Flex your Energy: Building a Cloud-Native Platform for Renewable Energy Communities](https://www.wearedevelopers.com/videos/1990-flex-your-energy-building-a-cloud-native-platform-for-renewable-energy-communities) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)