Senior Site Reliability Engineer (SRE) (Hybrid)
Role details
Job location
Tech stack
Job description
Experteer Overview In this role you will build and operate the Splunk Agent Observability deployment platform to ensure reliable, scalable AI agent deployments. You will own cross-environment deployments (cloud and air-gapped), drive deployment observability, and collaborate with engineering teams to deliver secure, resilient systems. You will automate for efficiency, participate in incident response, and tune performance to meet reliability goals. This hybrid role combines on-site work with Cisco offices in SF, San Jose, or NYC. Compensation / Benefits * Operate and improve Kubernetes-based production infrastructure and deployment systems * Own customer deployments across cloud and air-gapped environments, including installation, upgrades, troubleshooting, and lifecycle management * Build and improve deployment observability, monitoring, logging, and alerting * Improve reliability, scalability, and operational efficiency through automation and performance optimization * Participate in production incident response, root cause analysis, and reliability improvements * Tune infrastructure components-including databases and services-to improve performance and resilience * Design and develop internal tooling using Python and/or Go * Debug complex production issues spanning Kubernetes, networking, storage, and application layers * Manage infrastructure using Terraform or similar Infrastructure as Code tools * Collaborate with software engineers and customers to design secure, scalable, and reliable deployment architectures Tasks * 7+ years' experience with a Bachelor's degree or 4+ yrs with a Masters or 1 year with a PhD, or equivalent related experience * 3+ years' operating Kubernetes in production; experience with Helm * Experience building and maintaining CI/CD platforms and deployment automation * Experience working with AWS, GCP, or similar cloud platforms Key requirements * medical, dental, and vision insurance * 401(k) with matching contributions * paid parental leave * vacation and paid time off policies * stock grants potential * flexible vacation time off
Requirements
production incident response, root cause analysis, and reliability improvements * Tune infrastructure components-including databases and services-to improve performance and resilience * Design and develop internal tooling using Python and/or Go * Debug complex production issues spanning Kubernetes, networking, storage, and application layers * Manage infrastructure using Terraform or similar Infrastructure as Code tools * Collaborate with software engineers and customers to design secure, scalable, and reliable deployment architectures Tasks * 7+ years' experience with a Bachelor's degree or 4+ yrs with a Masters or 1 year with a PhD, or equivalent related experience * 3+ years' operating Kubernetes in production; experience with Helm * Experience building and maintaining CI/CD platforms and deployment automation * Experience working with AWS, GCP, or similar cloud platforms Key requirements * medical, dental, and vision insurance * 401(k) with matching contributions * paid parental leave * vacation and paid time off policies * stock grants potential * flexible vacation time off