> Markdown version of [/jobs/ext/2004477-staff-site-reliability-engineer-sre-hybrid](https://www.wearedevelopers.com/jobs/ext/2004477-staff-site-reliability-engineer-sre-hybrid). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Site Reliability Engineer (SRE) (Hybrid) - **Company:** Cisco Systems, Inc. - **Location:** New York, NY, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Databases, Continuous Integration, Distributed Systems, Reliability Engineering, Kubernetes, Deployment Automation, Splunk - **Published:** August 9, 2026 - **Apply:** https://us.experteer.com/career/view-jobs/staff-site-reliability-engineer-sre-hybrid-brooklyn-ny-usa-58863782 ## About the Role internal platforms and tooling for reliable operations at scale * Lead incident response and drive RCAs and long-term remediation * Collaborate with engineering leadership on platform architecture and production readiness * Mentor engineers and raise engineering standards through design reviews and guidelines * Work with customers and internal teams to design secure, scalable deployment architectures for cloud and on-prem environments Tasks * 8+ years' experience with a Bachelor's degree or 6+ years with a Masters or 3+ years with a PhD, or equivalent; at least 6 years in SRE/platform/cloud/infrastructure * 5+ years operating large-scale Kubernetes platforms in production * Experience designing highly available, scalable, resilient distributed systems * Experience with AWS, GCP, or other public cloud platforms * Strong experience designing CI/CD platforms and deployment automation at scale Key requirements * medical, dental, and vision insurance * 401(k) with matching * paid parental leave * short and long-term disability * basic life insurance * paid time away and holidays ## Description Experteer Overview In this role you will lead reliability for the Splunk Agent Observability platform, shaping the long-term strategy and owning large-scale cloud and on-prem deployments. You'll drive platform resilience, deployment automation, and incident response while mentoring engineers and guiding cross-team architecture. You'll partner with leadership to improve production readiness and scale reliability across environments. This is a hands-on leadership role that blends engineering excellence with strategic direction to support AI resilience at scale. Compensation / Benefits * Define the reliability roadmap for platform scalability and operational excellence * Specify and evolve deployment platforms for cloud and air-gapped environments * Set SLOs, readiness, capacity planning, and resiliency reviews * Lead reliability and scalability initiatives across Kubernetes, deployment infra, databases, and networking * Automate operations to reduce toil and boost productivity * Build internal platforms and tooling for reliable operations at scale * Lead incident response and drive RCAs and long-term remediation * Collaborate with engineering leadership on platform architecture and production readiness * Mentor engineers and raise engineering standards through design reviews and guidelines * Work with customers and internal teams to design secure, scalable deployment architectures for cloud and on-prem environments Tasks * 8+ years' experience with a Bachelor's degree or 6+ years with a Masters or 3+ years with a PhD, or equivalent; at least 6 years in SRE/platform/cloud/infrastructure * 5+ years operating large-scale Kubernetes platforms in production * Experience designing highly available, scalable, resilient distributed systems * Experience with AWS, GCP, or other public cloud platforms * Strong experience designing CI/CD platforms and deployment automation at scale Key requirements * medical, dental, and vision insurance * 401(k) with matching * paid parental leave * short and long-term disability * basic life insurance * paid time away and holidays ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Kubernetes and Microservices with Multi-Model Databases](https://www.wearedevelopers.com/videos/382-kubernetes-and-microservices-with-multi-model-databases) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/videos/100333-instant-kai-sandboxes-with-vcluster-multi-tenant-multi-scheduler-gpu-sharing) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [How Much Does a Software Engineer Make? Realistic Software Engineering Salaries](https://www.wearedevelopers.com/magazine/425-how-much-does-a-software-engineer-make-realistic-software-engineering-salaries) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers)