> Markdown version of [/jobs/ext/2728224-site-reliability-engineer-sre](https://www.wearedevelopers.com/jobs/ext/2728224-site-reliability-engineer-sre). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer (SRE) - **Company:** zuven Technologies - **Location:** Chandler, AZ, United States - **Experience:** Experienced - **Salary:** $125,299.0 - $141,939.0 - **Contract:** Temporary to permanent - **Skills:** Java (Programming Language), Artificial Intelligence, Amazon Web Services, Application Performance Management, Confluence, Automation of Tests, Microsoft Azure, BigQuery, Cloud Computing, Computer Programming, Databases, Continuous Integration, Software Debugging, DevOps, Distributed Systems, Domain Name System (DNS), Monitoring of Systems, Hypertext Transfer Protocols (HTTP), Python (Programming Language), PostgreSQL, Microsoft SQL Server, MongoDB, NoSQL, OpenShift, Oracle (Applications), Redis, Reliability Engineering, Prometheus, PL-SQL, TCP/IP, Google Cloud, Enterprise Software Applications, Load Balancing, Istio, Grafana, Generative AI, Containerization, Kubernetes, Rancher, Atlassian Tools, Hashicorp, Graphql, Cloud Migration, Vertica, Api Gateway, Splunk, Appdynamics, Dynatrace, Golang - **Published:** September 5, 2026 - **Apply:** https://www.careerjet.com/jobad/us3422aec319285efa1a19cfc263be06d2 ## About the Role We are seeking a Site Reliability Engineer with strong experience in cloud-native operations, observability, automation, and production support for large-scale enterprise applications. Skillset required: * 3-5 years of experience in Site Reliability Engineering, Production Operations, or Platform Engineering supporting large-scale, high-performance applications across hybrid environments (on-premises and cloud). * 3-5 years of experience developing automation scripts and building Application Performance Management (APM) dashboards to monitor end-to-end transaction journeys. * Hands-on programming experience (2+ years) with one or more languages such as Go, Python, Java, or Rust. * Working knowledge of relational and NoSQL databases including Oracle, SQL Server, PostgreSQL, MongoDB, Redis, ClickHouse, PL/SQL, or time-series databases. * Experience with cloud migration and containerization initiatives using GCP, AWS, Azure, Rancher, OpenShift, or similar platforms. * Experience managing containerized applications in Kubernetes environments such as GKE, RKE, or AKS. * Strong experience implementing observability solutions using Open Telemetry (OTEL), distributed tracing, monitoring, and incident management. * Familiarity with GraphQL frameworks such as Apollo, Prisma, or Hasura. * Strong networking fundamentals including TCP/IP, HTTP, DNS, load balancing, and service mesh technologies. * Experience participating in 24x7 on-call rotations and meeting incident response SLAs., * Experience managing highly available, customer-facing platforms with a focus on reliability, automation, and operational excellence. * Hands-on experience with monitoring and observability tools such as Splunk, Dynatrace, AppDynamics, Grafana, and Prometheus. * Experience with CI/CD and Agile tools such as Rally, Confluence, and related DevOps platforms. * Knowledge of in-memory caching technologies, especially Redis. * Strong troubleshooting and debugging skills across distributed systems and API gateway architectures. * Experience with Google Cloud services including GCS, Cloud SQL, Spanner, and BigQuery. * Experience supporting HashiCorp Vault environments. * Exposure to Vertex AI, Generative AI, and cloud-based analytics platforms. ## Description Job Details: Job Description: The Board-Level Quality and Reliability Engineer is responsible for ensuring the quality, robustness, and long-term reliability of advanced semico… + 30 days ago + ## Related Videos - [Leveraging Real time data in FSIs](https://www.wearedevelopers.com/videos/806-leveraging-real-time-data-in-fsis) - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Accelerating Authentication Architecture: Taking Passwordless to the Next Level](https://www.wearedevelopers.com/videos/733-accelerating-authentication-architecture-taking-passwordless-to-the-next-level) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs)