Staff Site Reliability Engineer (SRE) (Hybrid)

Cisco Systems, Inc.
New York, NY, United States
3 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Databases Continuous Integration Distributed Systems Reliability Engineering Kubernetes Deployment Automation Splunk

Job description

Experteer Overview In this role you will lead reliability for the Splunk Agent Observability platform, shaping the long-term strategy and owning large-scale cloud and on-prem deployments. You’ll drive platform resilience, deployment automation, and incident response while mentoring engineers and guiding cross-team architecture. You’ll partner with leadership to improve production readiness and scale reliability across environments. This is a hands-on leadership role that blends engineering excellence with strategic direction to support AI resilience at scale. Compensation / Benefits * Define the reliability roadmap for platform scalability and operational excellence * Specify and evolve deployment platforms for cloud and air-gapped environments * Set SLOs, readiness, capacity planning, and resiliency reviews * Lead reliability and scalability initiatives across Kubernetes, deployment infra, databases, and networking * Automate operations to reduce toil and boost productivity * Build internal platforms and tooling for reliable operations at scale * Lead incident response and drive RCAs and long-term remediation * Collaborate with engineering leadership on platform architecture and production readiness * Mentor engineers and raise engineering standards through design reviews and guidelines * Work with customers and internal teams to design secure, scalable deployment architectures for cloud and on-prem environments Tasks * 8+ years’ experience with a Bachelor’s degree or 6+ years with a Masters or 3+ years with a PhD, or equivalent; at least 6 years in SRE/platform/cloud/infrastructure * 5+ years operating large-scale Kubernetes platforms in production * Experience designing highly available, scalable, resilient distributed systems * Experience with AWS, GCP, or other public cloud platforms * Strong experience designing CI/CD platforms and deployment automation at scale Key requirements * medical, dental, and vision insurance * 401(k) with matching * paid parental leave * short and long-term disability * basic life insurance * paid time away and holidays

Requirements

internal platforms and tooling for reliable operations at scale * Lead incident response and drive RCAs and long-term remediation * Collaborate with engineering leadership on platform architecture and production readiness * Mentor engineers and raise engineering standards through design reviews and guidelines * Work with customers and internal teams to design secure, scalable deployment architectures for cloud and on-prem environments Tasks * 8+ years’ experience with a Bachelor’s degree or 6+ years with a Masters or 3+ years with a PhD, or equivalent; at least 6 years in SRE/platform/cloud/infrastructure * 5+ years operating large-scale Kubernetes platforms in production * Experience designing highly available, scalable, resilient distributed systems * Experience with AWS, GCP, or other public cloud platforms * Strong experience designing CI/CD platforms and deployment automation at scale Key requirements * medical, dental, and vision insurance * 401(k) with matching * paid parental leave * short and long-term disability * basic life insurance * paid time away and holidays

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · WWC 2023

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · WWC 2022

3:04 min

Database evolution and the funding behind vector databases

Erik Bamberg · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

4:04 min

Overview of Kubernetes operators and custom resource definitions

Philipp Krenn · WWC 2022

4:01 min

Managing application isolation via pluggable database models

Wei Hu Wei Hu · WWC 2022

Videos

See all

Related articles

See all