Senior Site Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+12 more
Job description
You will join our global team in Sunnyvale and help operate financial platforms at scale. You will partner with cross-functional teams to improve reliability, automate operations, and drive continuous improvement across our cloud-native environments.
What you will do:
- Design, build and maintain automation to eliminate manual, repetitive operational tasks (runbooks, deployment pipelines, remediation scripts).
- Operate and enhance monitoring, logging and alerting systems to ensure strong observability across services.
- Participate in on-call rotations and lead incident response activities; run and document post-incident RCA and follow-up actions.
- Collaborate with stakeholders to define SLIs and SLOs, manage error budgets and translate reliability goals into measurable actions.
- Forecast capacity needs and contribute to resource planning to ensure performance and cost-efficiency.
- Troubleshoot production issues: deep-dive analysis, isolate root causes and implement durable fixes.
- Drive continuous improvement and platform hardening through runbook improvements, automation, and best-practice adoption.
- Work closely as part of an international team to deliver project outcomes and operational excellence.
Requirements
- Solid practical experience in site reliability, operations or DevOps at a mid-to-senior level.
- Strong shell scripting skills and a foundation in programming concepts.
- Hands-on experience with cloud workloads-specifically Google Cloud Platform (GCP) and GKE.
- Proven experience with containerisation and orchestration (Kubernetes).
- Working knowledge of Infrastructure as Code and configuration management (Terraform, Ansible, Puppet).
- Familiarity with monitoring and observability tooling such as Prometheus, Grafana and Datadog.
- In-depth understanding of HTTP(s) traffic, routing and load-balancing, with practical experience observing and operating HAProxy.
- Comfortable using GitHub and GitHub Actions for code management, automation and IaC pipelines.
- Strong troubleshooting skills, a pragmatic problem-solving approach and effective communication for cross-team collaboration.
What would be great to have:
- Experience programming in Python, Go or Java.
- Exposure to large-scale financial services platforms or highly regulated environments.
- Experience defining SLIs/SLOs and managing error budgets in production environments.
About the company
We’re Fiserv, a global leader in Fintech and payments, and we move money and information in a way that moves the world. We connect financial institutions, corporations, merchants and consumers to one another millions of times a day - quickly, reliably, and securely. Any time you swipe your credit card, pay through a mobile app, or withdraw money from the bank, we’re involved. If you want to make an impact on a global scale, come make a difference at Fiserv.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Highest Paying Tech Companies for Developers
Is Software Engineering Over-Saturated?
How Much FAANG Companies Actually Pay Software Engineers in 2025
Fully Remote Software Engineer Jobs