Senior Site Reliability Engineer (Capacity) - Platform Infrastructure
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
OverviewAs Principal Platform Engineer focused on capacity, you optimize Elastic Cloud Hosted and Serverless workloads to scale seamlessly.You’ll partner with control plane and cross-functional teams to tackle cloud scaling challenges and resource allocation at a global scale.Your work directly supports customers by ensuring reliable, efficient compute across many regions.This role offers a meaningful impact on how Elastic delivers scalable AI-powered search and observability.Compensaciones / Beneficiosbase salaryhealth coverage for you and familyflexible locations and schedulesgenerous vacation daysdonations matching up to $**volunteering time (up to 40 hours/year) versus parity of benefits across regionsResponsabilidadesAssess current and future capacity needs and develop predictive modelsCollaborate to implement proactive capacity planning to avoid shortagesOptimize resource usage across cloud environments for performance and scalabilityAnalyze metrics and build reporting tools for visibility into capacity and performanceOperate an autoscaling framework for diverse customer workloadsOptimize infrastructure performance across 60+ Elastic Cloud regionsWork with development teams to implement scaling best practicesRequisitos principales5+ years in cloud infrastructure and capacity managementKnowledge of performance monitoring and optimization techniquesUnderstanding of cloud scaling challenges and solutionsProficiency in incident investigation and troubleshooting processesExperience with compute auto-scaling and capacity reservations across the three major CSPsSolid software and platform engineering backgroundExperience with all three major cloud service providers and navigating compute capacity scaling issuesproblem-solving mindsetcross-functional collaborationproactive communicationcloud infrastructure managementcapacity planning and modelingperformance monitoring and tuning
Requirements
Knowledge of performance monitoring and optimization techniques Understanding of cloud scaling challenges and solutions Proficiency in incident investigation and troubleshooting processes Experience with compute auto-scaling and capacity reservations across the three major CSPs Solid software and platform engineering background Experience with all three major cloud service providers and navigating compute capacity scaling issues problem-solving mindset cross-functional collaboration proactive communication cloud infrastructure management capacity planning and modeling performance monitoring and tuning
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
What Are The Top Skills Required For Azure Developers?
7 Cloud Computing Trends Coming in 2025 for Developers
Is Software Engineering Over-Saturated?
Fully Remote Software Engineer Jobs