Site Reliability Engineer

Tieto
Madrid, Spain
13 days ago
Apply on www.buscojobs.com.es
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
Chinese, English

Tech stack

Batch Processing Business Software Cloud Computing Cloud Engineering Continuous Delivery DevOps Middleware Monitoring of Systems Python (Programming Language) Enterprise Messaging Systems Nginx Octopus Deploy
+12 more
RabbitMQ Reliability Engineering Prometheus TCP/IP Virtual Machines Scripting Delivery Pipeline Grafana Containerization Apache Kafka Ci Server Microservices

Job description

About the PositionWe are looking for anSRE Consultantto join an international, cloud-native environment, responsible for the operation, stability, automation, and evolution of cloud infrastructures and business applications deployed across Europe.The selected candidate will work closely with technical and operations teams based in Spain and China, contributing to architecture standardization, process automation, observability, and the continuous improvement of service performance, availability, and efficiency.ResponsibilitiesManage the end-to-endoperation, maintenance, and architectural governanceof public cloud infrastructures in Europe, including containers, virtual machines, storage, and networking.Ensure the operational stability and architectural standardization ofdatabases and middleware.Design, implement, and maintainautomated CI/CD pipelines, enabling standardized integration, delivery, and continuous deployment processes.Participate inon-call rotations, providing support for incidents and production issues.Proactively identify, diagnose, and resolvefunctional incidents, resource bottlenecks, and performance anomalies.Manage the daily operation, maintenance, andproduction releasesof cloud-native and containerized applications.Ensure the availability, stability, and proper evolution of production services.Collaborate with technical teams in China to define and implementcommon management standards for containerized applications.Drive initiatives focused onstandardization, technology governance, architecture optimization, and application migrationtowards cloud-native environments.Continuously analyze cloud and application resource consumption and utilization, identifying opportunities forcapacity, performance, and cost optimization.Promote continuous improvement in the efficiency and utilization of resources across European environments.Ensure that operational and maintenance activities comply withsecurity, privacy, and applicable regulatory requirements in Europe.Required Skills and ExperienceStrong experience inLinux administration and operations, with an in-depth understanding of its core principles and components.Advanced knowledge ofnetworking and protocols, particularly TCP/IP andHands-on experience withKubernetes environments, including a solid understanding of the architecture and operation of its main components.Experience operating and maintainingKubernetes Operators in production environments.Experience withobservability and monitoring tools, particularlyPrometheus and Grafana, with the ability to design and maintain comprehensive monitoring systems for cloud-native environments.Experience withCI/CD tools and methodologies, particularlyArgoCD, and the ability to build, configure, and maintain automated pipelines.Knowledge of the architecture, deployment, and operation ofAPI Gateways, particularly:NginxAPISIXEnvoyKnowledge ofmiddleware and messaging systems, particularly:RocketMQRabbitMQKafkaProficiency in at least one scripting language, preferablyPython or Shell, for automation, batch processing, and the development of operational tools.Knowledge ofGDPR, privacy, and data protection regulations applicable in Europe.Ability to apply regulatory compliance requirements to infrastructure and cloud operations.Nice-to-Have SkillsExperience inmulticloud or public cloud environments.Experience withcloud-native architectures and microservices.Experience in operations automation andInfrastructure as Code (IaC)practices.Knowledge of cloud cost optimization andFinOps.Experience working with international and geographically distributed technical teams.Experience in high-availability environments and business-critical systems.Experience with application migration processes towards cloud-native architectures.Knowledge ofSRE, DevOps, and GitOpsmethodologies.Experience in incident management,Root Cause Analysis (RCA), and continuous improvement.LanguagesEnglish:Advanced level, essential for communication with international teams.Chinese:Highly desirable due to ongoing collaboration with engineering and operations teams based in China.Candidates withthe legal right to work in Spainand the ability to operate effectively in international environments will be particularly valued.Soft SkillsStrong analytical and problem-solving skills.Strong focus on service stability, availability, and quality.Automation mindset and commitment to continuous improvement.Ability to work under pressure and manage production incidents.Excellent communication and collaboration skills when working with multidisciplinary technical teams.Ability to work independently and takeend-to-end ownershipof services.Strong focus on standardization, efficiency, and resource optimization.Adaptability to working in an international and multicultural environment.What We OfferThe opportunity to join ahighly technology-driven, international cloud-native environment.Participation in European-wide infrastructure and application projects.Direct collaboration with international technology teams.The opportunity to work with technologies such asKubernetes, Prometheus, Grafana, ArgoCD, Nginx, APISIX, Envoy, Kafka, RabbitMQ, and Python/Shell.Participation in cloud architecture automation, standardization, optimization, and evolution projects.Professional development opportunities within a dynamic and rapidly growing technology environment.LocationSpain, with regular collaboration with technical and operations teams based in China.

Requirements

Strong experience inLinux administration and operations, with an in-depth understanding of its core principles and components. Advanced knowledge ofnetworking and protocols, particularly TCP/IP and Hands-on experience withKubernetes environments, including a solid understanding of the architecture and operation of its main components. Experience operating and maintainingKubernetes Operators in production environments. Experience withobservability and monitoring tools, particularlyPrometheus and Grafana, with the ability to design and maintain comprehensive monitoring systems for cloud-native environments. Experience withCI/CD tools and methodologies, particularlyArgoCD, and the ability to build, configure, and maintain automated pipelines. Knowledge of the architecture, deployment, and operation ofAPI Gateways, particularly: Nginx APISIX Envoy Knowledge ofmiddleware and messaging systems, particularly: RocketMQ RabbitMQ Kafka Proficiency in at least one scripting language, preferablyPython or Shell, for automation, batch processing, and the development of operational tools. Knowledge ofGDPR, privacy, and data protection regulations applicable in Europe. Ability to apply regulatory compliance requirements to infrastructure and cloud operations. Nice-to-Have Skills Experience inmulticloud or public cloud environments. Experience withcloud-native architectures and microservices. Experience in operations automation andInfrastructure as Code (IaC)practices. Knowledge of cloud cost optimization andFinOps. Experience working with international and geographically distributed technical teams. Experience in high-availability environments and business-critical systems. Experience with application migration processes towards cloud-native architectures. Knowledge ofSRE, DevOps, and GitOpsmethodologies. Experience in incident management,Root Cause Analysis (RCA), and continuous improvement. Languages English:Advanced level, essential for communication with international teams. Chinese:Highly desirable due to ongoing collaboration with engineering and operations teams based in China. Candidates withthe legal right to work in Spainand the ability to operate effectively in international environments will be particularly valued. Soft Skills Strong analytical and problem-solving skills. Strong focus on service stability, availability, and quality. Automation mindset and commitment to continuous improvement. Ability to work under pressure and manage production incidents. Excellent communication and collaboration skills when working with multidisciplinary technical teams. Ability to work independently and takeend-to-end ownershipof services. Strong focus on standardization, efficiency, and resource optimization. Adaptability to working in an international and multicultural environment.

About the company

The opportunity to work with technologies such asKubernetes, Prometheus, Grafana, ArgoCD, Nginx, APISIX, Envoy, Kafka, RabbitMQ, and Python/Shell. Participation in cloud architecture automation, standardization, optimization, and evolution projects. Professional development opportunities within a dynamic and rapidly growing technology environment. Location Spain, with regular collaboration with technical and operations teams based in China.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.buscojobs.com.es
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:06 min

Developer experience and project variety at scale

Alexandra Petri · World Congress 2023

5:02 min

Mapping distributed compute paradigms to modern vehicles

Joachim Werner · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

7:28 min

Constructing a new Docker layer from scratch

Oliver Seitz Oliver Seitz · World Congress 2026 Europe

3:50 min

Queues in TCP stacks and continuous network connections

Clemens Vasters Clemens Vasters · World Congress 2022

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all