> Markdown version of [/jobs/ext/2596685-senior-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2596685-senior-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer - **Company:** Fiserv, Inc. - **Location:** Sunnyvale, CA, United States - **Experience:** Expert - **Salary:** $160,000.0 - $240,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Cloud Computing, Configuration Management, DevOps, Github, HAProxy, HTTP Secure, Python (Programming Language), Routing, Reliability Engineering, Cloud Services, Ansible, Prometheus, Shell Script, Datadog, Data Logging, Load Balancing, Delivery Pipeline, Grafana, Containerization, Kubernetes, Puppet, Terraform, Software Version Control - **Published:** August 19, 2026 - **Apply:** https://dejobs.org/x/x/8BDD38A141884AD1B2A98DB1AAD62122/job/ ## About the Role * Solid practical experience in site reliability, operations or DevOps at a mid-to-senior level. * Strong shell scripting skills and a foundation in programming concepts. * Hands-on experience with cloud workloads-specifically Google Cloud Platform (GCP) and GKE. * Proven experience with containerisation and orchestration (Kubernetes). * Working knowledge of Infrastructure as Code and configuration management (Terraform, Ansible, Puppet). * Familiarity with monitoring and observability tooling such as Prometheus, Grafana and Datadog. * In-depth understanding of HTTP(s) traffic, routing and load-balancing, with practical experience observing and operating HAProxy. * Comfortable using GitHub and GitHub Actions for code management, automation and IaC pipelines. * Strong troubleshooting skills, a pragmatic problem-solving approach and effective communication for cross-team collaboration. What would be great to have: * Experience programming in Python, Go or Java. * Exposure to large-scale financial services platforms or highly regulated environments. * Experience defining SLIs/SLOs and managing error budgets in production environments. ## Description You will join our global team in Sunnyvale and help operate financial platforms at scale. You will partner with cross-functional teams to improve reliability, automate operations, and drive continuous improvement across our cloud-native environments. What you will do: * Design, build and maintain automation to eliminate manual, repetitive operational tasks (runbooks, deployment pipelines, remediation scripts). * Operate and enhance monitoring, logging and alerting systems to ensure strong observability across services. * Participate in on-call rotations and lead incident response activities; run and document post-incident RCA and follow-up actions. * Collaborate with stakeholders to define SLIs and SLOs, manage error budgets and translate reliability goals into measurable actions. * Forecast capacity needs and contribute to resource planning to ensure performance and cost-efficiency. * Troubleshoot production issues: deep-dive analysis, isolate root causes and implement durable fixes. * Drive continuous improvement and platform hardening through runbook improvements, automation, and best-practice adoption. * Work closely as part of an international team to deliver project outcomes and operational excellence. ## Related Videos - [Inside Bitpanda's Tech Stack: Scaling a European Fintech Leader - Markus Dorner](https://www.wearedevelopers.com/videos/1979-inside-bitpanda-s-tech-stack-scaling-a-european-fintech-leader-markus-dorner) - [Automate everything via NodeJS and Puppeteer](https://www.wearedevelopers.com/videos/322-automate-everything-via-nodejs-and-puppeteer) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [The Memory Leak That Ate Our Cluster: A Postmortem](https://www.wearedevelopers.com/videos/2057-the-memory-leak-that-ate-our-cluster-a-postmortem) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [How Much FAANG Companies Actually Pay Software Engineers in 2025](https://www.wearedevelopers.com/magazine/230-how-much-faang-companies-actually-pay-software-engineers-in-2025) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)