Senior Site Reliability Engineer (Sre)

Epam Systems
Málaga, Spain
18 days ago
Apply on www.buscojobs.com.es
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Cloud Computing Continuous Integration DevOps Monitoring of Systems Python (Programming Language) Machine Learning Systems Development Life Cycle Reliability Engineering Site Reliability Engineering Practices
+12 more
Ansible Scripting Cloud Platform System System Availability Large Language Models Gitlab Containerization Kubernetes Information Technology Terraform Docker Jenkins

Job description

We’re looking for aSenior Site Reliability Engineer (SRE)to join our team in Spain in a remote working mode. In this role, you will collaborate with development, operations, security and quality teams to ensure highly reliable, scalable and efficient systems for business-critical applications in the financial domain. You will focus on implementing SRE practices, reducing toil through automation and driving operational excellence while meeting strict Service Level Objectives (SLOs).This position offers the opportunity to influence system design for reliability and performance within a global delivery context, leveraging modern cloud technologies, observability tools and automation frameworks to maintain seamless user experiences.Define and maintain Service Level Objectives (SLOs), SLIs and error budgets for critical services Collaborate with cross-functional teams to embed reliability into application and infrastructure design Automate operational tasks to reduce manual toil and improve service performance Troubleshoot and resolve infrastructure and application incidents quickly and effectively Implement robust monitoring and observability systems to detect and prevent outages Plan capacity and scaling strategies to ensure high availability and resiliency Contribute to incident postmortems and continuous improvement initiatives Support the adoption of SRE best practices across all SDLC stages Bachelor’s degree in Computer Science, Engineering or related field Proven experience working in cloud environments (AWS, GCP or Azure) Practical knowledge of SRE principles (SLO/SLI design, error budgets, postmortems, automation) Proficiency in Python or other scripting language for automation tasks Strong understanding of monitoring tools and observability frameworks Experience with Infrastructure-as-Code and CI/CD tools (e.g., Terraform, Ansible, Jenkins, GitLab) Hands-on expertise with containerization and orchestration platforms such as Docker and Kubernetes Experience deploying and managing Large Language Models (LLMs), including RAG-based solutions Certifications in Kubernetes, AWS/GCP/Azure or related cloud technologies Background in DevOps practices and agile delivery frameworks Familiarity with AI/ML model operations: deployment, monitoring and optimization in production environments

Requirements

This position offers the opportunity to influence system design for reliability and performance within a global delivery context, leveraging modern cloud technologies, observability tools and automation frameworks to maintain seamless user experiences. Define and maintain Service Level Objectives (SLOs), SLIs and error budgets for critical services Collaborate with cross-functional teams to embed reliability into application and infrastructure design Automate operational tasks to reduce manual toil and improve service performance Troubleshoot and resolve infrastructure and application incidents quickly and effectively Implement robust monitoring and observability systems to detect and prevent outages Plan capacity and scaling strategies to ensure high availability and resiliency Contribute to incident postmortems and continuous improvement initiatives Support the adoption of SRE best practices across all SDLC stages Bachelor’s degree in Computer Science, Engineering or related field Proven experience working in cloud environments (AWS, GCP or Azure) Practical knowledge of SRE principles (SLO/SLI design, error budgets, postmortems, automation) Proficiency in Python or other scripting language for automation tasks Strong understanding of monitoring tools and observability frameworks Experience with Infrastructure-as-Code and CI/CD tools (e.g., Terraform, Ansible, Jenkins, GitLab) Hands-on expertise with containerization and orchestration platforms such as Docker and Kubernetes Experience deploying and managing Large Language Models (LLMs), including RAG-based solutions Certifications in Kubernetes, AWS/GCP/Azure or related cloud technologies Background in DevOps practices and agile delivery frameworks Familiarity with AI/ML model operations: deployment, monitoring and optimization in production environments

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.buscojobs.com.es
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · World Congress 2021

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

6:14 min

Structuring CI/CD pipelines with integrated security and quality checks

Christoph Ruggenthaler · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

1:38 min

Adopting site reliability engineering practices for machine learning

Cassie Kozyrkov · World Congress 2022

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

Videos

See all

Related articles

See all