Senior Site Reliability Engineer (Sre)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+12 more
Job description
We’re looking for aSenior Site Reliability Engineer (SRE)to join our team in Spain in a remote working mode. In this role, you will collaborate with development, operations, security and quality teams to ensure highly reliable, scalable and efficient systems for business-critical applications in the financial domain. You will focus on implementing SRE practices, reducing toil through automation and driving operational excellence while meeting strict Service Level Objectives (SLOs).This position offers the opportunity to influence system design for reliability and performance within a global delivery context, leveraging modern cloud technologies, observability tools and automation frameworks to maintain seamless user experiences.Define and maintain Service Level Objectives (SLOs), SLIs and error budgets for critical services Collaborate with cross-functional teams to embed reliability into application and infrastructure design Automate operational tasks to reduce manual toil and improve service performance Troubleshoot and resolve infrastructure and application incidents quickly and effectively Implement robust monitoring and observability systems to detect and prevent outages Plan capacity and scaling strategies to ensure high availability and resiliency Contribute to incident postmortems and continuous improvement initiatives Support the adoption of SRE best practices across all SDLC stages Bachelor’s degree in Computer Science, Engineering or related field Proven experience working in cloud environments (AWS, GCP or Azure) Practical knowledge of SRE principles (SLO/SLI design, error budgets, postmortems, automation) Proficiency in Python or other scripting language for automation tasks Strong understanding of monitoring tools and observability frameworks Experience with Infrastructure-as-Code and CI/CD tools (e.g., Terraform, Ansible, Jenkins, GitLab) Hands-on expertise with containerization and orchestration platforms such as Docker and Kubernetes Experience deploying and managing Large Language Models (LLMs), including RAG-based solutions Certifications in Kubernetes, AWS/GCP/Azure or related cloud technologies Background in DevOps practices and agile delivery frameworks Familiarity with AI/ML model operations: deployment, monitoring and optimization in production environments
Requirements
This position offers the opportunity to influence system design for reliability and performance within a global delivery context, leveraging modern cloud technologies, observability tools and automation frameworks to maintain seamless user experiences. Define and maintain Service Level Objectives (SLOs), SLIs and error budgets for critical services Collaborate with cross-functional teams to embed reliability into application and infrastructure design Automate operational tasks to reduce manual toil and improve service performance Troubleshoot and resolve infrastructure and application incidents quickly and effectively Implement robust monitoring and observability systems to detect and prevent outages Plan capacity and scaling strategies to ensure high availability and resiliency Contribute to incident postmortems and continuous improvement initiatives Support the adoption of SRE best practices across all SDLC stages Bachelor’s degree in Computer Science, Engineering or related field Proven experience working in cloud environments (AWS, GCP or Azure) Practical knowledge of SRE principles (SLO/SLI design, error budgets, postmortems, automation) Proficiency in Python or other scripting language for automation tasks Strong understanding of monitoring tools and observability frameworks Experience with Infrastructure-as-Code and CI/CD tools (e.g., Terraform, Ansible, Jenkins, GitLab) Hands-on expertise with containerization and orchestration platforms such as Docker and Kubernetes Experience deploying and managing Large Language Models (LLMs), including RAG-based solutions Certifications in Kubernetes, AWS/GCP/Azure or related cloud technologies Background in DevOps practices and agile delivery frameworks Familiarity with AI/ML model operations: deployment, monitoring and optimization in production environments
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Is Software Engineering Over-Saturated?
Where To Find Software Engineering Jobs
Find a Developer Job: 12 Best Job Sites For Developers
Why Upskilling And Reskilling is Important For Developers