Site Reliability Engineer

Valid
Madrid, Spain
11 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
3 years minimum
Working hours
Shift work
Languages
English

Tech stack

Amazon Web Services Backup Devices Bash Shell Ubuntu (Operating System) Cloud Computing Program Optimization Databases Computer Engineering Continuous Integration Linux DevOps Disaster Recovery
+23 more
Github Monitoring of Systems Python (Programming Language) Linux System Administration Performance Tuning Red Hat Enterprise Linux Reliability Engineering Ansible Prometheus Standard Sql Datadog System Availability Grafana Reliability of Systems Containerization Gitlab-ci Kubernetes Digital Government Performance Monitor Terraform Docker Elk Stack Jenkins

Job description

If you’re passionate about technology, innovative projects, and making a real impact, your place is here.Cualquier información adicional que necesite para este trabajo se encuentra en el texto a continuación.Asegúrese de leerla detenidamente y luego envíe su solicitud.We are a global technology provider with 65+ years of experience, delivering a comprehensive portfolio of solutions across ID & Digital Government, Banking & Payments, and Trusted Connectivity.With more than 4,000 employees in 16 countries, we are committed to building a more secure and trustworthy world.Within ourTrusted Connectivity business unit , we develop cutting?edge solutions for the telecommunications industry-ranging fromSIM cards and eSIMs to Subscription Management and secure connectivity services -connecting people, businesses, and devices worldwide.We are looking for a highly analytical and business?orientedSite Reliability Engineerto design and enhance the reliability engineering architecture of our platforms, ensuring high availability, scalability, reliability, and observability through close collaboration with R&D, DevOps, and Operations teams.As aSite Reliability Engineer , you will be responsible for working mainly with Site Reliability Architect (SRA) and R&D team to design resilient systems and operational processes that ensure the high availability, scalability, reliability and observability of our platforms.What will you do?Site Reliability EngineeringWork together with R&D to develop and maintain reliable, scalable, and efficient systems.Work closely with R&D when new features are being developed and ensure that the new feature is ready to be released.Ensure new features have been validated in terms of performance, reliability and scalability.Prepare and conduct knowledge transfer, documentation and information sharing to the other team members.Cross-functional CollaborationWork together with DevOps team to improve existing and implement new, effective CI/CD processes.Work together with Enablement engineer to produce automation tools needed for performance and reliability monitoring.Work together with Operations team to support the platforms in terms of operational aspects.Continuously evaluate and optimize system performance and capacity in order to maintain stable production platforms.Identify, assess, and implement measures to eliminate potential risks that could impact the performance of systems and services.Research, evaluate, test and advise at selecting appropriate new technologies or tools for improving site reliability.Observability & MonitoringMonitor system performance, identifying bottlenecks, and execute pipeline optimization.Implement comprehensive service metrics to track and report on system reliability, performance, and efficiency.Disaster Recovery & BackupsImplement disaster recovery plans and ensuring robust backup systems are in place.Capacity Planning & Performance EngineeringSupport in forecasting, scaling, and performance tuning.Create KPI to monitor growth and optimize resource utilization.MigrationsAnalyze and plan for complex migrations.What are we looking for?Bachelor’s degree in Computer Engineering, Electronics Engineering, Telecommunications Engineering, or a related field.3+ years of experience in Site Reliability Engineering, Infrastructure Operations, DevOps, or a similar role.3+ years of experience within the telecommunications industry or related technology sectors.Strong Linux administration skills (Red Hat, Ubuntu, or similar distributions).Hands?on experience with cloud platforms,preferably AWS.Experience designing and maintaining CI/CD pipelines (Jenkins, GitLab CI, GitHub Actions, or similar).Experience with monitoring and observability tools (Prometheus, Grafana, ELK Stack, Datadog, or equivalent).Proficiency in scripting and automation using Python, Bash, or similar languages.Experience with Infrastructure as Code (Terraform, Ansible, or equivalent).Strong knowledge of containerization and orchestration technologies (Docker, Kubernetes).Experience in performance monitoring, troubleshooting, and system optimization.Knowledge of disaster recovery, backup strategies, and business continuity practices.Experience working with SQL databases.Advanced English communication skills (B2+/C1).Candidates must be based in Spain or nearby European countries and be available to travel when required.If you want this position to be yours, we would like you to have the following:AWS, Linux, or Kubernetes certifications.Experience in highly available and mission?critical environments.Knowledge of capacity planning and performance engineering.Experience in telecom platforms, mobile services, or cloud?native architectures.What we offerJoin Valid and work on innovative, global technology projects within multicultural and multidisciplinary teams.Flexibility: flexible working hours and remote work options to support work?life balance.Well?being first: private medical insurance and life insurance.Be part of a company that values continuous learning, collaboration, and growth.Our CultureAt Valid, we foster an inclusive, diverse, and innovative environment where everyone thrives.We are committed to equal opportunities, free from discrimination concerning sex, age, race, sexual orientation, religion, education, social status, culture, or special needs such as illness or disability.We value people as the heart of our culture.Trust, transparency, and teamwork are the foundations of our success, driving growth and empowering talent.xkdbapoJoin this great team and be part of our story!#J-*****-Ljbffr

Requirements

Bachelor’s degree in Computer Engineering, Electronics Engineering, Telecommunications Engineering, or a related field. 3+ years of experience in Site Reliability Engineering, Infrastructure Operations, DevOps, or a similar role. 3+ years of experience within the telecommunications industry or related technology sectors. Strong Linux administration skills (Red Hat, Ubuntu, or similar distributions). Hands?on experience with cloud platforms, preferably AWS. Experience designing and maintaining CI/CD pipelines (Jenkins, GitLab CI, GitHub Actions, or similar). Experience with monitoring and observability tools (Prometheus, Grafana, ELK Stack, Datadog, or equivalent). Proficiency in scripting and automation using Python, Bash, or similar languages. Experience with Infrastructure as Code (Terraform, Ansible, or equivalent). Strong knowledge of containerization and orchestration technologies (Docker, Kubernetes). Experience in performance monitoring, troubleshooting, and system optimization. Knowledge of disaster recovery, backup strategies, and business continuity practices. Experience working with SQL databases. Advanced English communication skills (B2+/C1). Candidates must be based in Spain or nearby European countries and be available to travel when required. If you want this position to be yours, we would like you to have the following: AWS, Linux, or Kubernetes certifications. Experience in highly available and mission?critical environments. Knowledge of capacity planning and performance engineering. Experience in telecom platforms, mobile services, or cloud?native architectures.

Benefits & conditions

Join Valid and work on innovative, global technology projects within multicultural and multidisciplinary teams. Flexibility: flexible working hours and remote work options to support work?life balance. Well?being first: private medical insurance and life insurance. Be part of a company that values continuous learning, collaboration, and growth. Our Culture At Valid, we foster an inclusive, diverse, and innovative environment where everyone thrives. We are committed to equal opportunities, free from discrimination concerning sex, age, race, sexual orientation, religion, education, social status, culture, or special needs such as illness or disability. We value people as the heart of our culture. Trust, transparency, and teamwork are the foundations of our success, driving growth and empowering talent. xkdbapo Join this great team and be part of our story! #J-*****-Ljbffr

About the company

Madrid, España

If you’re passionate about technology, innovative projects, and making a real impact, your place is here. Cualquier información adicional que necesite para este trabajo se encuentra en el texto a continuación. Asegúrese de leerla detenidamente y luego envíe su solicitud. We are a global technology provider with 65+ years of experience, delivering a comprehensive portfolio of solutions across ID & Digital Government, Banking & Payments, and Trusted Connectivity. With more than 4,000 employees in 16 countries, we are committed to building a more secure and trustworthy world. Within our Trusted Connectivity business unit , we develop cutting?edge solutions for the telecommunications industry-ranging from SIM cards and eSIMs to Subscription Management and secure connectivity services -connecting people, businesses, and devices worldwide. We are looking for a highly analytical and business?oriented

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.buscojobs.com.es

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · WWC 2023

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

2:39 min

Experiencing core Linux capabilities for DevOps administration

Michael Cade · LIVE

8:02 min

Integrating service level objectives into incident management

Diana Todea · LIVE

Videos

See all

Related articles

See all