> Markdown version of [/jobs/ext/2080228-senior-site-reliability-engineer-sre](https://www.wearedevelopers.com/jobs/ext/2080228-senior-site-reliability-engineer-sre). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer (Sre) - **Company:** Epam Systems - **Location:** Málaga, Spain (Remote available) - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Microsoft Azure, Cloud Computing, Continuous Integration, DevOps, Monitoring of Systems, Python (Programming Language), Machine Learning, Systems Development Life Cycle, Reliability Engineering, Site Reliability Engineering Practices, Ansible, Scripting, Cloud Platform System, System Availability, Large Language Models, Gitlab, Containerization, Kubernetes, Information Technology, Terraform, Docker, Jenkins - **Published:** August 16, 2026 - **Apply:** https://www.buscojobs.com.es/senior-site-reliability-engineer-sre-en-malaga-ID-367454410 ## About the Role This position offers the opportunity to influence system design for reliability and performance within a global delivery context, leveraging modern cloud technologies, observability tools and automation frameworks to maintain seamless user experiences. Define and maintain Service Level Objectives (SLOs), SLIs and error budgets for critical services Collaborate with cross-functional teams to embed reliability into application and infrastructure design Automate operational tasks to reduce manual toil and improve service performance Troubleshoot and resolve infrastructure and application incidents quickly and effectively Implement robust monitoring and observability systems to detect and prevent outages Plan capacity and scaling strategies to ensure high availability and resiliency Contribute to incident postmortems and continuous improvement initiatives Support the adoption of SRE best practices across all SDLC stages Bachelor's degree in Computer Science, Engineering or related field Proven experience working in cloud environments (AWS, GCP or Azure) Practical knowledge of SRE principles (SLO/SLI design, error budgets, postmortems, automation) Proficiency in Python or other scripting language for automation tasks Strong understanding of monitoring tools and observability frameworks Experience with Infrastructure-as-Code and CI/CD tools (e.g., Terraform, Ansible, Jenkins, GitLab) Hands-on expertise with containerization and orchestration platforms such as Docker and Kubernetes Experience deploying and managing Large Language Models (LLMs), including RAG-based solutions Certifications in Kubernetes, AWS/GCP/Azure or related cloud technologies Background in DevOps practices and agile delivery frameworks Familiarity with AI/ML model operations: deployment, monitoring and optimization in production environments ## Description We're looking for aSenior Site Reliability Engineer (SRE)to join our team in Spain in a remote working mode. In this role, you will collaborate with development, operations, security and quality teams to ensure highly reliable, scalable and efficient systems for business-critical applications in the financial domain. You will focus on implementing SRE practices, reducing toil through automation and driving operational excellence while meeting strict Service Level Objectives (SLOs).This position offers the opportunity to influence system design for reliability and performance within a global delivery context, leveraging modern cloud technologies, observability tools and automation frameworks to maintain seamless user experiences.Define and maintain Service Level Objectives (SLOs), SLIs and error budgets for critical services Collaborate with cross-functional teams to embed reliability into application and infrastructure design Automate operational tasks to reduce manual toil and improve service performance Troubleshoot and resolve infrastructure and application incidents quickly and effectively Implement robust monitoring and observability systems to detect and prevent outages Plan capacity and scaling strategies to ensure high availability and resiliency Contribute to incident postmortems and continuous improvement initiatives Support the adoption of SRE best practices across all SDLC stages Bachelor's degree in Computer Science, Engineering or related field Proven experience working in cloud environments (AWS, GCP or Azure) Practical knowledge of SRE principles (SLO/SLI design, error budgets, postmortems, automation) Proficiency in Python or other scripting language for automation tasks Strong understanding of monitoring tools and observability frameworks Experience with Infrastructure-as-Code and CI/CD tools (e.g., Terraform, Ansible, Jenkins, GitLab) Hands-on expertise with containerization and orchestration platforms such as Docker and Kubernetes Experience deploying and managing Large Language Models (LLMs), including RAG-based solutions Certifications in Kubernetes, AWS/GCP/Azure or related cloud technologies Background in DevOps practices and agile delivery frameworks Familiarity with AI/ML model operations: deployment, monitoring and optimization in production environments ## Related Videos - [WeAreDevelopers LIVE - Modern DevOps for IoT Devices and More](https://www.wearedevelopers.com/videos/1805-wearedevelopers-live-modern-devops-for-iot-devices-and-more) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [The 12 Best Jobs for Software Engineers](https://www.wearedevelopers.com/magazine/401-the-12-best-jobs-for-software-engineers)