> Markdown version of [/jobs/ext/1917833-senior-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1917833-senior-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer - **Company:** DEVELOCITY GROUP, INC. - **Location:** San Francisco, CA, United States (Remote available) - **Experience:** Expert - **Salary:** $150,000.0 - $190,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, Backup Devices, Bash Shell, Software as a Service, Cloud Computing, DevOps, Disaster Recovery, Java Virtual Machine (JVM), Python (Programming Language), Open Source Technology, Reliability Engineering, Site Reliability Engineering Practices, Prometheus, Software Deployment, Data Logging, Scripting, Cloud Platform System, Grafana, Kotlin, Amazon Relational Database Service, Kubernetes, Asynchronous Programming, Terraform - **Published:** August 4, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=1fbb235d39ff0220 ## About the Role * 5+ years in SRE, DevOps, or equivalent role operating production services at scale. * Strong Kubernetes experience in production environments. * Cloud infrastructure expertise, preferably AWS (EKS, RDS, S3, EC2). * Proficiency with observability tools (Prometheus, Grafana) and Infrastructure as Code (Terraform). * Track record of incident management and response. * Knowledge of SRE best practices (SLAs, SLOs). * Scripting proficiency (Python, Bash) for automation. * Experience with 24/7 on-call rotations. * Strong written and verbal English communication., * Experience operating SaaS platforms at scale. * Familiarity with Develocity. * JVM language experience (Java, Kotlin). * Disaster recovery planning and execution experience. * Customer-facing incident communication skills. * Experience establishing SRE practices in new or growing teams. ## Description We're building a new SRE team and looking for founding members to help shape how we operate. You'll be responsible for the reliability, performance, and availability of Develocity instances serving paying customers, open-source projects, and public-facing services, plus supporting infrastructure like artifact registries. You'll work on our internally-built Cloud Application Platform, Kubernetes on AWS, and develop deep expertise in it. When incidents happen, you'll troubleshoot issues across the stack, from application to infrastructure. You'll collaborate with the Cloud Platform team to improve the tooling you depend on, and with engineering teams to build reliability into how we ship software. If you like automating things and hate doing the same task twice, you'll fit in well. You'll be part of a distributed, remote-first team that values asynchronous communication and written documentation. Strong self-direction and clear communication across time zones are essential., * Operate and maintain all Develocity instances and supporting services. * Participate in a follow-the-sun on-call rotation, owning incident response and troubleshooting issues across the stack. * Drive automation across application deployment, upgrades, monitoring, self-healing, and recovery. * Build and maintain observability for all managed services (logging, metrics, tracing, and alerting). * Work with engineering teams to build reliability into features from the start. * Run incident response and retrospectives, and make sure we learn from them. * Own disaster recovery, backups, and business continuity. * Communicate with customers during incidents and maintenance windows. * Optimize performance, resource usage, and costs. * Help evolve our SaaS operations as we grow. ## Related Videos - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Kotlin Multiplatform - True power of native code reuse](https://www.wearedevelopers.com/videos/4-kotlin-multiplatform-true-power-of-native-code-reuse) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [React Developer Salary [2023]](https://www.wearedevelopers.com/magazine/198-react-developer-salary-2023) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Best Countries for Software Engineers](https://www.wearedevelopers.com/magazine/267-best-countries-for-software-engineers)