Staff Site Reliability Engineer (SRE) (Hybrid)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
Experteer Overview In this role you will lead reliability for the Splunk Agent Observability platform, shaping the long-term strategy and owning large-scale cloud and on-prem deployments. You’ll drive platform resilience, deployment automation, and incident response while mentoring engineers and guiding cross-team architecture. You’ll partner with leadership to improve production readiness and scale reliability across environments. This is a hands-on leadership role that blends engineering excellence with strategic direction to support AI resilience at scale. Compensation / Benefits * Define the reliability roadmap for platform scalability and operational excellence * Specify and evolve deployment platforms for cloud and air-gapped environments * Set SLOs, readiness, capacity planning, and resiliency reviews * Lead reliability and scalability initiatives across Kubernetes, deployment infra, databases, and networking * Automate operations to reduce toil and boost productivity * Build internal platforms and tooling for reliable operations at scale * Lead incident response and drive RCAs and long-term remediation * Collaborate with engineering leadership on platform architecture and production readiness * Mentor engineers and raise engineering standards through design reviews and guidelines * Work with customers and internal teams to design secure, scalable deployment architectures for cloud and on-prem environments Tasks * 8+ years’ experience with a Bachelor’s degree or 6+ years with a Masters or 3+ years with a PhD, or equivalent; at least 6 years in SRE/platform/cloud/infrastructure * 5+ years operating large-scale Kubernetes platforms in production * Experience designing highly available, scalable, resilient distributed systems * Experience with AWS, GCP, or other public cloud platforms * Strong experience designing CI/CD platforms and deployment automation at scale Key requirements * medical, dental, and vision insurance * 401(k) with matching * paid parental leave * short and long-term disability * basic life insurance * paid time away and holidays
Requirements
internal platforms and tooling for reliable operations at scale * Lead incident response and drive RCAs and long-term remediation * Collaborate with engineering leadership on platform architecture and production readiness * Mentor engineers and raise engineering standards through design reviews and guidelines * Work with customers and internal teams to design secure, scalable deployment architectures for cloud and on-prem environments Tasks * 8+ years’ experience with a Bachelor’s degree or 6+ years with a Masters or 3+ years with a PhD, or equivalent; at least 6 years in SRE/platform/cloud/infrastructure * 5+ years operating large-scale Kubernetes platforms in production * Experience designing highly available, scalable, resilient distributed systems * Experience with AWS, GCP, or other public cloud platforms * Strong experience designing CI/CD platforms and deployment automation at scale Key requirements * medical, dental, and vision insurance * 401(k) with matching * paid parental leave * short and long-term disability * basic life insurance * paid time away and holidays
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on us.experteer.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Is Software Engineering Over-Saturated?
Fully Remote Software Engineer Jobs
How Much Does a Software Engineer Make? Realistic Software Engineering Salaries
Dev Digest 120 - Apple and peers