Devops/Mlops Engineer (Re3)

Barcelona Supercomputing Center
Madrid, Spain
7 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
3 years minimum
Working hours
Regular working hours
Languages
English

Tech stack

Artificial Intelligence Airflow Cloud Computing Configuration Management Information Systems DevOps Octopus Deploy Performance Tuning Release Management Reliability Engineering Runbook Scientific Computating
+16 more
Software Deployment Software Engineering Supercomputing Management of Software Versions Data Logging Large Language Models Software Troubleshooting Generative AI Infrastructure as Code (IaC) Containerization AI Platforms Kubernetes Infrastructure Automation Frameworks Information Technology Machine Learning Operations Docker

Job description

About BSC The Barcelona Supercomputing Center - Centro Nacional de Supercomputación (BSC?CNS) is the leading supercomputing centre in Spain.It houses MareNostrum, one of Europe’s most powerful supercomputers, and hosts EuroHPC JU.Its mission is to research, develop, and manage information technologies to facilitate scientific progress.BSC combines HPC service provision and R&D in computer and computational science under one roof and has over ** staff from 60 countries.PositionDevOps/MLOps Engineer (RE3) - Reference 331_26_DIR_IBD_RE3.Full?time, 35 h/week, located at BSC within the Directors Department.Closing date: Friday, 31 July **.Key DutiesDesign, implement, and maintain CI/CD pipelines for AI and machine?learning solutions.Support deployment, versioning, release management, and rollback processes for AI services.Implement monitoring, logging, alerting, and observability capabilities for AI systems and infrastructure.Manage model registries, experiment?tracking platforms, and deployment metadata.Automate infrastructure provisioning, configuration management, and operational workflows.Collaborate with AI engineers, platform teams, security teams, and cloud administrators.Support environment standardisation, reproducibility, and operational best practices.Define operational procedures, runbooks, and incident?response processes.Support troubleshooting, performance optimisation, and continuous improvement of production AI services.Contribute to MLOps, DevOps, and platform?engineering standards across the AI Factory.RequirementsEducationBachelor’s or Master’s degree in Computer Science, Software Engineering, Artificial Intelligence, Telecommunications Engineering, Information Systems, or a related technical discipline.Essential Knowledge and Professional Experience3-5 years of experience in DevOps, MLOps, Platform Engineering, Site Reliability Engineering (SRE), or related roles.Experience designing and operating CI/CD pipelines in production environments.Experience with containerisation technologies such as Docker and orchestration platforms such as Kubernetes.Knowledge of Infrastructure as Code (IaC) practices and automation frameworks.Experience with monitoring, logging, observability, and alerting solutions.Understanding of machine?learning and generative?AI deployment lifecycle requirements.Experience with model registries, experiment?tracking tools, and release?management processes.Knowledge of security, access control, environment management, and operational governance.Strong troubleshooting, automation, and operational problem?solving skills.Additional Knowledge and Professional ExperienceExperience supporting LLM, RAG, or generative?AI platforms.Familiarity with MLflow, Kubeflow, Airflow, ArgoCD, or similar MLOps ecosystems.Experience in HPC, scientific computing, or large?scale AI infrastructures.Experience supporting multi?tenant AI platforms and shared?services environments.Experience working across infrastructure, AI, security, and platform teams.Adaptability in a fast?evolving AI and cloud technology environment.Strong DevOps, MLOps, and automation skills.CompetencesAbility to build reliable, scalable, and maintainable operational environments.Strong analytical and troubleshooting capabilities.Proactive, autonomous, and results?oriented mindset.Strong focus on reliability, observability, and operational excellence.Commitment to continuous improvement and automation.ConditionsContract: Open?ended full?time (35 h/week) within the Directors Department.Workplace: BSC, flexible working hours, extensive training plan, restaurant tickets, private health insurance, and relocation support.Holidays: 22 days + 6 personal days + the 24th and 31st of December per collective agreement.Competitive salary commensurate with qualifications and the cost of living in Barcelona.Starting date: 16 July **.Application ProcedureApplicants must submit a CV in English and a cover letter in English via the BSC website.Two references may be required.All documents should be named using the structureName_Surname_CV ,Name_Surname_CoverLetter , etc., and must be under the permitted file sizes.Equal Opportunity StatementBSC?CNS is an equal?opportunity employer committed to diversity and inclusion.We consider all qualified applicants for employment without regard to race, colour, religion, sex, sexual orientation, gender identity, national origin, age, disability or any other basis protected by applicable state or local law.#J-***-Ljbffr

Requirements

Bachelor’s or Master’s degree in Computer Science, Software Engineering, Artificial Intelligence, Telecommunications Engineering, Information Systems, or a related technical discipline. Essential Knowledge and Professional Experience 3-5 years of experience in DevOps, MLOps, Platform Engineering, Site Reliability Engineering (SRE), or related roles. Experience designing and operating CI/CD pipelines in production environments. Experience with containerisation technologies such as Docker and orchestration platforms such as Kubernetes. Knowledge of Infrastructure as Code (IaC) practices and automation frameworks. Experience with monitoring, logging, observability, and alerting solutions. Understanding of machine?learning and generative?AI deployment lifecycle requirements. Experience with model registries, experiment?tracking tools, and release?management processes. Knowledge of security, access control, environment management, and operational governance. Strong troubleshooting, automation, and operational problem?solving skills. Additional Knowledge and Professional Experience Experience supporting LLM, RAG, or generative?AI platforms. Familiarity with MLflow, Kubeflow, Airflow, ArgoCD, or similar MLOps ecosystems. Experience in HPC, scientific computing, or large?scale AI infrastructures. Experience supporting multi?tenant AI platforms and shared?services environments. Experience working across infrastructure, AI, security, and platform teams. Adaptability in a fast?evolving AI and cloud technology environment. Strong DevOps, MLOps, and automation skills. Competences Ability to build reliable, scalable, and maintainable operational environments. Strong analytical and troubleshooting capabilities. Proactive, autonomous, and results?oriented mindset. Strong focus on reliability, observability, and operational excellence. Commitment to continuous improvement and automation. Conditions

Benefits & conditions

Contract: Open?ended full?time (35 h/week) within the Directors Department. Workplace: BSC, flexible working hours, extensive training plan, restaurant tickets, private health insurance, and relocation support. Holidays: 22 days + 6 personal days + the 24th and 31st of December per collective agreement. Competitive salary commensurate with qualifications and the cost of living in Barcelona. Starting date: 16 July **.

About the company

Madrid, España

About BSC The Barcelona Supercomputing Center - Centro Nacional de Supercomputación (BSC?CNS) is the leading supercomputing centre in Spain. It houses MareNostrum, one of Europe’s most powerful supercomputers, and hosts EuroHPC JU. Its mission is to research, develop, and manage information technologies to facilitate scientific progress. BSC combines HPC service provision and R&D in computer and computational science under one roof and has over ** staff from 60 countries.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.buscojobs.com.es

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · WWC Europe 2026

2:44 min

Defining core roles and responsibilities in MLOps teams

Bas Geerdink · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · WWC Europe 2026

3:05 min

Audience questions on AI agents and pipeline vectorization

Joy Joy · WWC 2024

Videos

See all

Related articles

See all