Senior Site Reliability Engineer (Capacity) - Platform Infrastructure

Elastic
Burgos, Spain
12 days ago
Apply on www.buscojobs.com.es
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Amazon Elastic Compute Cloud Cloud Computing Network Control Reliability Engineering Cloud Services Autoscaling Optimization Algorithms Performance Monitor Tools for Reporting Serverless Computing

Job description

OverviewAs Principal Platform Engineer focused on capacity, you optimize Elastic Cloud Hosted and Serverless workloads to scale seamlessly.You’ll partner with control plane and cross-functional teams to tackle cloud scaling challenges and resource allocation at a global scale.Your work directly supports customers by ensuring reliable, efficient compute across many regions.This role offers a meaningful impact on how Elastic delivers scalable AI-powered search and observability.Compensaciones / Beneficiosbase salaryhealth coverage for you and familyflexible locations and schedulesgenerous vacation daysdonations matching up to $**volunteering time (up to 40 hours/year) versus parity of benefits across regionsResponsabilidadesAssess current and future capacity needs and develop predictive modelsCollaborate to implement proactive capacity planning to avoid shortagesOptimize resource usage across cloud environments for performance and scalabilityAnalyze metrics and build reporting tools for visibility into capacity and performanceOperate an autoscaling framework for diverse customer workloadsOptimize infrastructure performance across 60+ Elastic Cloud regionsWork with development teams to implement scaling best practicesRequisitos principales5+ years in cloud infrastructure and capacity managementKnowledge of performance monitoring and optimization techniquesUnderstanding of cloud scaling challenges and solutionsProficiency in incident investigation and troubleshooting processesExperience with compute auto-scaling and capacity reservations across the three major CSPsSolid software and platform engineering backgroundExperience with all three major cloud service providers and navigating compute capacity scaling issuesproblem-solving mindsetcross-functional collaborationproactive communicationcloud infrastructure managementcapacity planning and modelingperformance monitoring and tuning

Requirements

Knowledge of performance monitoring and optimization techniques Understanding of cloud scaling challenges and solutions Proficiency in incident investigation and troubleshooting processes Experience with compute auto-scaling and capacity reservations across the three major CSPs Solid software and platform engineering background Experience with all three major cloud service providers and navigating compute capacity scaling issues problem-solving mindset cross-functional collaboration proactive communication cloud infrastructure management capacity planning and modeling performance monitoring and tuning

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.buscojobs.com.es
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:39 min

Modernizing operational excellence and cloud deployment tools

Mustafa Toroman · World Congress 2023

3:33 min

Managing node capacity with default cluster autoscaling

Mario-Leander Reimer · World Congress 2023

2:41 min

Transitioning artificial intelligence infrastructure into scalable commodity cloud services

juarezjunior juarezjunior · World Congress 2024

2:27 min

Introduction to WebAssembly in a cloud computing context

Edo Edo · World Congress 2024

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:48 min

Leveraging compute separation for autoscaling and Git-like branching

Oleksandra Bovkun Oleksandra Bovkun · Europe 2026 Virtual

Videos

See all

Related articles

See all