Site Reliability Engineer

ProntoPro
Spain
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Application Performance Management Microsoft Azure Cloud Computing Program Optimization Computer Programming Fault Tolerance Github Monitoring of Systems Python (Programming Language) Nagios Ansible
+13 more
Software Engineering Systems Architecture Datadog System Availability Multi-Cloud Reliability of Systems Infrastructure as Code (IaC) Cloudformation Kubernetes Infrastructure Automation Frameworks Cloudwatch Terraform Docker

Job description

CloudFormation for consistent cloud environment provisioning - Design and deploy containerized solutions using Docker and Kubernetes on cloud platforms Cloud Infrastructure & Operations - Participate in architectural decisions and implement cloud infrastructure solutions using Azure and AWS services, ensuring high availability and scalability - Manage and optimize cloud resources to improve performance, cost efficiency, and security - Integrate advanced monitoring and alerting tools (e.g., Datadog, CloudWatch, Application Insights) to maintain system observability in multi-cloud environments Incident Management & Analysis - Lead incident response, conduct root cause analyses, and produce blameless postmortems to prevent future occurrences - Collaborate with development teams to integrate observability and performance metrics into the development lifecycle - Build and maintain executive-level and developer-centric dashboards to visualize key metrics Technical Expertise - Proven

Requirements

experience in running and maintaining production systems with expertise in triaging and solving incidents - Proficiency in automation and configuration management tools (e.g., Terraform, Ansible) - Expertise in cloud platforms, particularly Azure and AWS, and their associated tools - Strong programming skills, with a primary focus on Python, for developing automation scripts, creating custom tooling, and optimizing operational workflows - Experience with modern observability platforms such as Datadog Skills & Experience - A solid foundation in system architecture, with a focus on scalability and reliability - Exceptional problem-solving skills and a data-driven mindset Desirable Requirements - Experience with container orchestration tools such as Kubernetes - Familiarity with CI/CD pipelines and tools like GitHub Actions and Azure DevOps - Knowledge of security best practices in cloud and hybrid environments J-18808-Ljbffr

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on es.trabajo.org

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:37 min

Why differing legacy workflows complicate monitoring tool migrations

Mathias Palmersheim Mathias Palmersheim · Europe 2026 Virtual

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · WWC 2023

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:40 min

Using GitHub primitives for internal documentation and corporate operations

Kyle Daigle · Coffee With Developers

2:19 min

Applying code assistant capabilities to infrastructure and cloud operations

Ryan J Salva · Coffee With Developers

Videos

See all

Related articles

See all