> Markdown version of [/jobs/ext/2561379-senior-site-reliability-engineer-production-engineer-thousandeyes](https://www.wearedevelopers.com/jobs/ext/2561379-senior-site-reliability-engineer-production-engineer-thousandeyes). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer, Production Engineer - ThousandEyes - **Company:** Cisco Systems, Inc. - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Salary:** $165,000.0 - $241,400.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Client Server Models, Software as a Service, File Systems, Distributed Systems, Python (Programming Language), Octopus Deploy, Reliability Engineering, Prometheus, Software Engineering, Istio, Reliability of Systems, Kubernetes, Low Latency, Deployment Automation - **Published:** August 1, 2026 - **Apply:** https://jobs.localjobnetwork.com/apply/add/87910421/1 ## About the Role * 5+ years of experience in a related role * Proficiency in software development with languages such as Python or Go * Shown ability to build and implement scalable, well-tested, and security-focused solutions that integrate security protocols throughout the development and deployment lifecycle * Strong understanding of Unix/Linux systems, including kernel, system libraries, file systems, and client-server protocols * Knowledge of Site Reliability principles: Incident Response, Change Management, Distributed Systems, Deployment Strategies, and SLOs, * Familiarity with procedures for operating a large-scale, highly available enterprise platform * Excellent communication and documentation skills * Strong sense of ownership, drive, and attention to detail * Expert-level knowledge of Kubernetes and its ecosystem * In-depth knowledge of cloud providers, preferably AWS ## Description We are seeking a skilled Senior Site Reliability Engineer (SRE) in Production Engineering with a strong background in SaaS and operations. You will design and manage large-scale, highly available distributed systems in the cloud, collaborating directly with application development teams to enhance the reliability, performance, and security of our platform., * Collaborate with software engineers to optimize architecture and services for availability, latency, performance, and reliability using cloud-native tools. * Design and implement scalable operations tooling to support platform growth and scaling across multiple regions. * Design, deploy, and maintain AWS cloud-native services that are elastic and resilient to failure. * Participate in and improve our 24x7 incident response and on-call rotation. * Use and expand our existing CNCF solutions like Kubernetes, Service Mesh, Prometheus, OpenTelemetry, and ArgoCD to increase platform reliability. * Automate production operations to provide guardrails and continuous platform operation. * Develop automation solutions for scalable service and platform operations, including deployment, scale testing, graceful failure, and chaos testing. * Stay updated on industry best practices for scalability and reliability to improve the scalability of the ThousandEyes platform. * Identify and provide solutions to common obstacles hindering operational excellence across engineering teams. * Generalize and standardize solutions and processes to enable repeated success across our microservice-based multi-region platform. * Play a key role in the ThousandEyes platform by leveraging scale testing, additional environments, and working with application teams to improve system reliability. * Manage a rapidly growing infrastructure capable of handling substantial daily data volumes, emphasizing operations/infrastructure/everything as code. ## Related Videos - [How Cisco embraced a DevOps culture within its network engineering team](https://www.wearedevelopers.com/videos/99-how-cisco-embraced-a-devops-culture-within-its-network-engineering-team) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Data binning and understanding histograms](https://www.wearedevelopers.com/videos/2086-data-binning-and-understanding-histograms) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [Microservices: how to get started with Spring Boot and Kubernetes](https://www.wearedevelopers.com/videos/242-microservices-how-to-get-started-with-spring-boot-and-kubernetes) - [Next-gen CI/CD with Gitops and Progressive Delivery](https://www.wearedevelopers.com/videos/1603-next-gen-ci-cd-with-gitops-and-progressive-delivery) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [The Best Software Developer Blogs to Read](https://www.wearedevelopers.com/magazine/156-the-best-software-developer-blogs-to-read) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence)