Staff Site Reliability Engineer (SRE) (Hybrid)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
Experteer Overview As Staff Site Reliability Engineer, you will lead reliability, scalability, and operational architecture for Splunk Agent Observability’s platform. You will set long-term reliability strategy, drive major infrastructure initiatives, and raise engineering standards across deployment automation, production operations, and platform resiliency. You’ll influence platform architecture and guide cloud and on-prem deployments, collaborating with cross-functional teams and customers. This role offers impact at scale, shaping how AI-enabled deployments are observed, controlled, and trusted. You’ll partner with leadership to elevate incident response and drive long-term improvements. Compensation / Benefits * Define and drive the technical roadmap for platform reliability, scalability and operational excellence * Lead architecture and evolution of deployment platforms for cloud and air-gapped environments * Establish reliability standards including SLOs, readiness, capacity planning, and resiliency reviews * Drive major reliability and scalability initiatives across Kubernetes, deployment infra, databases, and networking * Automate to reduce toil and boost engineering productivity * Build internal platforms and tooling for reliable operations at scale * Lead incident response and postmortems with long-term remediation * Partner with engineering leadership on platform architecture and deployment strategy * Mentor engineers via design reviews and operational guidelines * Collaborate with customers and teams to design secure, scalable deployment architectures for cloud and on-prem environments Tasks * 8+ years’ experience with a Bachelor’s degree or 6+ yrs with Masters or 3+ years with a PhD, or equivalent related experience * At least 6 years in Site Reliability Engineering, Platform/Cloud/Infrastructure Engineering, or related fields * 5+ years operating large-scale Kubernetes platforms in production * Experience designing highly available, scalable distributed systems * Experience with AWS, GCP, or other public clouds * Strong experience designing CI/CD platforms and deployment automation at scale Key requirements * medical, dental and vision insurance * 401(k) with matching contribution * paid parental leave * paid holidays and vacation policies * sick time and personal wellness days * volunteer days (optional)
Requirements
with and resiliency reviews * Drive major reliability and scalability initiatives across Kubernetes, deployment infra, databases, and networking * Automate to reduce toil and boost engineering productivity * Build internal platforms and tooling for reliable operations at scale * Lead incident response and postmortems with long-term remediation * Partner with engineering leadership on platform architecture and deployment strategy * Mentor engineers via design reviews and operational guidelines * Collaborate with customers and teams to design secure, scalable deployment architectures for cloud and on-prem environments Tasks * 8+ years’ experience with a Bachelor’s degree or 6+ yrs with Masters or 3+ years with a PhD, or equivalent related experience * At least 6 years in Site Reliability Engineering, Platform/Cloud/Infrastructure Engineering, or related fields * 5+ years operating large-scale Kubernetes platforms in production * Experience designing highly available, scalable distributed systems * Experience with AWS, GCP, or other public clouds * Strong experience designing CI/CD platforms and deployment automation at scale Key requirements * medical, dental and vision insurance * 401(k) with matching contribution * paid parental leave * paid holidays and vacation policies * sick time and personal wellness days * volunteer days (optional)
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on us.experteer.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Fully Remote Software Engineer Jobs
Highest Paying Tech Companies for Developers
Dev Digest 120 - Apple and peers
Find a Developer Job: 12 Best Job Sites For Developers