Senior Site Reliability Engineer - Hybrid & Full Time (Eu/Uk Only) (Castro)

Comply365
Municipality of Burgos, Spain
3 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Shift work
Languages
English
Experience level
Senior

Job location

Municipality of Burgos, Spain

Tech stack

Artificial Intelligence
Amazon Web Services (AWS)
Software as a Service
Cloud Computing
Cloud Engineering
Continuous Integration
Linux
Distributed Systems
Fault Tolerance
Python
Operational Data Store
Reliability Engineering
Site Reliability Engineering Practices
Prometheus
Systems Architecture
Datadog
Data Logging
Pulumi
Data Processing
Cloud Platform System
System Availability
Grafana
Infrastructure Automation Frameworks
Information Technology
Deployment Automation
Terraform
Dynatrace
Microservices

Job description

ph3Background /h3pHow would you feel about shaping the future of aviation safety while leading our move from legacy infrastructure to AWS?/ppAt SafetyManager365, we move quickly, build with intent, and focus on solving problems that genuinely matter to our customers and the aviation industry.Airlines generate huge volumes of safety reports, operational data and investigations.Our platform helps safety teams analyse that information, identify risks earlier and make better operational decisions.As a newly formed team within the wider Comply365 group, we bring the energy and mindset of a startup, with the backing and stability of an established enterprise.Our AI and data processing workloads already run entirely on AWS, and we are now extending that approach as we modernise the rest of the platform.We are looking for a Senior Site Reliability Engineer to play a key role in this journey - enhancing reliability, shaping our cloud architecture, reducing operational toil, and mentoring our infrastructure team as they adopt modern cloud and SRE practices.This role is for someone who is genuinely energised by moving things forward - who spots a manual process and instinctively asks how it could be automated, and who pushes for durable solutions ("one-and-done") rather than repeating the same fixes.We are at an early stage of our cloud maturity journey and need an individual who is comfortable dealing with ambiguity, not someone who reaches for a checklist./ppThis is a high impact role that sits at the intersection of reliability engineering, cloud infrastructure, and engineering enablement, and will be adecuado for someone who wants to combine technical leadership with hands?on engineering.You will play a key role in defining how we build, operate, and scale the platform.You will partner closely with software engineers to improve production readiness, deployment safety, observability, and operational maturity across the platform, reducing future incidents through better system design, simplification, and stronger operational foundations.The role is full?time (40 hours per week) with core hours from 10am to 5pm CET to ensure strong team alignment and collaboration, while still allowing flexibility outside of those hours.It is hybrid with at least 2 days in our offices in Barcelona per week and reports into the technical leadership within the SafetyManager365 group./ph3Key Responsibilities /h3ulliProactively identify risks, weaknesses, and operational bottlenecks, introducing structured solutions to move the platform from reactive to reliable and strategic /liliTake end?to?end ownership of reliability and infrastructure challenges, driving them through to durable resolution by treating root causes, not symptoms, and keeping standards above short?term workarounds /liliDesign, build, and continuously improve AWS infrastructure with a focus on scalability, resilience, performance, and cost efficiency /liliLead the migration of services and workloads from legacy dedicated environments into modern, cloud?native AWS architectures /liliDevelop and maintain Infrastructure as Code to automate provisioning, reduce manual effort, and ensure consistency across environments.Proficiency in at least one scripting or general?purpose language (Python preferred) is required; you will be asked to work through a practical coding or automation problem as part of the interview process /liliEnhance CI/CD pipelines to enable safe, fast, and repeatable deployments, improving overall developer productivity and system stability /liliMentor and support infrastructure engineers, promoting modern cloud and SRE practices, and leveraging AI tools where appropriate to improve automation, efficiency, and engineering outcomes /liliIdentify and eliminate toil: where a task is recurring and automatable, automate it.Apply judgment to prioritise where automation delivers the most leverage, while documenting areas that aren't yet, thereby avoiding single?points of failure./liliBuild and evolve observability capabilities across metrics, logging, tracing, and alerting to enable effective monitoring and rapid incident response /liliEstablish best practices for incident management and collaborate with engineering teams to improve production readiness, resilience, performance, and system architecture /li /ulh3Skills Qualifications /h3ulliAt least 8 years of experience in this or similar roles.Significant experience running production SaaS platforms in a senior SRE, platform, or infrastructure engineering role is required./liliA genuinely proactive mindset - not just responsive, but anticipatory.You notice problems before they escalate, propose concrete solutions without being asked, and stay uncomfortable with the status quo.If you are content to do repetitive manual work rather than drive to automate it, this role is likely not for you./liliProven ability to operate independently and think strategically in complex, ambiguous environments, using sound judgement and strong problem?solving skills, without waiting for the whole picture./liliDeep, hands?on expertise with AWS, including designing and operating scalable, resilient, and cost?efficient cloud architectures./liliDemonstrated experience migrating and modernising legacy infrastructure into cloud?native environments./liliStrong foundations in Linux, networking, and systems internals, with experience designing for high availability, fault tolerance, and production resilience at scale./liliProficiency with Infrastructure as Code tools such as Terraform, Pulumi, or similar, alongside experience improving CI/CD and automation practices./liliStrong communication and collaboration skills, with a mentoring mindset coupled with experience of working closely with engineering teams to improve reliability, observability, and operational maturity./li /ulh3Nice to haves /h3ulliDeep expertise in AWS, with a proven ability to architect, operate, and optimise secure, scalable, resilient, and cost?efficient cloud platforms; AWS certifications are advantageous /liliProven track record defining and managing SLIs, SLOs, and error budgets, supported by hands?on experience with modern observability platforms (e.g., Prometheus, Grafana, Datadog) and distributed tracing /liliExperience operating distributed systems and microservices architectures, including approaches such as service meshes, chaos engineering, and resilience testing /liliExposure to AI/ML or large?scale data workloads in production, with experience working in regulated or compliance?sensitive environments and a pragmatic approach to leveraging AI to improve engineering outcomes /li /ulh3Benefits /h3ulliEquipment: Laptop (Macbook), monitor, and whatever else you need to get productive /liliAnnual learning budget: courses, books, conferences, coaching /liliConference and speaking budget: attend industry events /liliAI tooling budget: pick the tools that make you better, we will cover them /liliAnnual team offsite /liliSalary range: 100,****,000 Euros annually /li /ulpYour application implies your consent to the processing of your personal data as outlined in our Privacy Policy./p /p #J--Ljbffr

Requirements

liliBuild and evolve observability capabilities across metrics, logging, tracing, and alerting to enable effective monitoring and rapid incident response /liliEstablish best practices for incident management and collaborate with engineering teams to improve production readiness, resilience, performance, and system architecture /li /ulh3Skills Qualifications /h3ulliAt least 8 years of experience in this or similar roles. Significant experience running production SaaS platforms in a senior SRE, platform, or infrastructure engineering role is required. /liliA genuinely proactive mindset - not just responsive, but anticipatory. You notice problems before they escalate, propose concrete solutions without being asked, and stay uncomfortable with the status quo. If you are content to do repetitive manual work rather than drive to automate it, this role is likely not for you. /liliProven ability to operate independently and think strategically in complex, ambiguous environments, using sound judgement and strong problem?solving skills, without waiting for the whole picture. /liliDeep, hands?on expertise with AWS, including designing and operating scalable, resilient, and cost?efficient cloud architectures. /liliDemonstrated experience migrating and modernising legacy infrastructure into cloud?native environments. /liliStrong foundations in Linux, networking, and systems internals, with experience designing for high availability, fault tolerance, and production resilience at scale. /liliProficiency with Infrastructure as Code tools such as Terraform, Pulumi, or similar, alongside experience improving CI/CD and automation practices. /liliStrong communication and collaboration skills, with a mentoring mindset coupled with experience of working closely with engineering teams to improve reliability, observability, and operational maturity. /li /ulh3Nice to haves /h3ulliDeep expertise in AWS, with a proven ability to architect, operate, and optimise secure, scalable, resilient, and cost?efficient cloud platforms; AWS certifications are advantageous /liliProven track record defining and managing SLIs, SLOs, and error budgets, supported by hands?on experience with modern observability platforms (e.g., Prometheus, Grafana, Datadog) and distributed tracing /liliExperience operating distributed systems and microservices architectures, including approaches such as service meshes, chaos engineering, and resilience testing /liliExposure to AI/ML or large?scale data workloads in production, with experience working in regulated or compliance?sensitive environments and a pragmatic approach to leveraging AI to improve engineering outcomes /li /ulh3Benefits /h3ulliEquipment: Laptop (Macbook), monitor, and whatever else you need to get productive /liliAnnual learning budget: courses, books, conferences, coaching /liliConference and speaking budget: attend industry events /liliAI tooling budget: pick the tools that make you better, we will cover them /liliAnnual team offsite /liliSalary range: 100,*********,000 Euros annually /li /ulpYour application implies your consent to the processing of your personal data as outlined in our Privacy Policy.

About the company

ph3Background /h3pHow would you feel about shaping the future of aviation safety while leading our move from legacy infrastructure to AWS?

Apply for this position