Ai & Hpc Workload Orchestration Engineer

Roche
Madrid, Spain
8 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
5 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Systems Engineering Distributed Systems InfiniBand Node.Js Red Hat Enterprise Linux System Availability Multi-Agent Systems Containerization Kubernetes Infrastructure Automation Frameworks Information Technology
+1 more
Slurm

Job description

At Roche you can show up as yourself, embraced for the unique qualities you bring.Our culture encourages personal expression, open dialogue, and genuine connections, where you are valued, accepted and respected for who you are, allowing you to thrive both personally and professionally.This is how we aim to prevent, stop and cure diseases and ensure everyone has access to healthcare today and for generations to come.Join Roche, where every voice matters.Job ResponsibilitiesOrchestration Stack Deployment & GovernanceDesign, implement, and maintain the SLURM Workload Manager ecosystem across HPC cluster architectures, ensuring high availability and optimal resource distribution.Deploy and manage Run:ai as the core orchestration and virtualization layer for the AI Factory, enabling fractional GPU allocation and dynamic resource allocation.Evaluate, architect, and implement SLURM Slinky integrations where required to seamlessly bridge Kubernetes-based AI orchestration with traditional HPC cluster resources.Containerization & Workload OptimizationDefine best practices and frameworks for containerized scientific execution, utilizing Singularity/Apptainer and/or Enroot to provide secure, reproducible performance environments for HPC.Translate user and workload requirements into optimized scheduling parameters (e.g., topology-aware scheduling, multi-node scaling).Actively profile and tune scheduling queues, QoS parameters, and fair-share policies to maximize multi-tenant efficiency.Platform Reliability & TelemetryPartner with Observability Engineers to implement continuous monitoring, telemetry, and reporting dashboards to track scheduler efficiency, queue wait times, and hardware utilization rates.Troubleshoot complex workload failures, including distributed training synchronization issues, MPI communication bottlenecks, and driver incompatibilities.Maintain configuration-as-code models for the scheduling tier, leveraging automation to deploy cluster policies uniformly.QualificationsEducation / Experience : Bachelor’s or advanced degree in Computer Science, Applied Mathematics, Computational Engineering, or a similar technical discipline.5+ years of systems engineering experience, with a heavy emphasis on workload scheduling, resource management, and cluster optimization for multi-tenant environments.Technical Proficiency : Deep familiarity with Enterprise Linux operating systems and distributed systems architecture.Expert-level proficiency in administering SLURM, including complex partition designs, accounting, and plug-in management.Highly proficient with Singularity for container runtime execution.AI Orchestration : Hands?on experience or deep architectural understanding of Run:ai, Kubernetes, and containerized GPU scheduling paradigms.Infrastructure Literacy : Solid understanding of high?speed interconnects (InfiniBand, RoCE) and multi?node communication architectures (MPI, NCCL) as they relate to job placement.Automation : Proficiency in automating scheduler configurations and telemetry gathering, or infrastructure automation tooling.Leadership & Mindset : Lean & Agile mindset; highly focused on driving efficiency, reducing idle compute time, and creating frictionless pathways for user workload submissions.Collaboration & Advocacy: Outstanding capability to translate scientific and AI model workflow challenges into scalable scheduler configurations.Intellectual Curiosity: A strong passion for remaining ahead of industry trends regarding GPU slicing, fractionalization, and the convergence of AI workloads with traditional HPC schedulers.Roche is an Equal Opportunity Employer.#J-*****-Ljbffr

Requirements

Education / Experience : Bachelor’s or advanced degree in Computer Science, Applied Mathematics, Computational Engineering, or a similar technical discipline. 5+ years of systems engineering experience, with a heavy emphasis on workload scheduling, resource management, and cluster optimization for multi-tenant environments. Technical Proficiency : Deep familiarity with Enterprise Linux operating systems and distributed systems architecture. Expert-level proficiency in administering SLURM, including complex partition designs, accounting, and plug-in management. Highly proficient with Singularity for container runtime execution. AI Orchestration : Hands?on experience or deep architectural understanding of Run:ai, Kubernetes, and containerized GPU scheduling paradigms. Infrastructure Literacy : Solid understanding of high?speed interconnects (InfiniBand, RoCE) and multi?node communication architectures (MPI, NCCL) as they relate to job placement. Automation : Proficiency in automating scheduler configurations and telemetry gathering, or infrastructure automation tooling. Leadership & Mindset : Lean & Agile mindset; highly focused on driving efficiency, reducing idle compute time, and creating frictionless pathways for user workload submissions. Collaboration & Advocacy: Outstanding capability to translate scientific and AI model workflow challenges into scalable scheduler configurations. Intellectual Curiosity: A strong passion for remaining ahead of industry trends regarding GPU slicing, fractionalization, and the convergence of AI workloads with traditional HPC schedulers.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.buscojobs.com.es

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:36 min

Managing new AI workloads for non-technical employees

Michael Coté Michael Coté · WWC Europe 2026

1:57 min

Routing cross-rack traffic seamlessly with NCCL

Kevin Klues Kevin Klues · WWC 2025

45 sec

Working securely with Node.js path application programming interfaces

Sonya Moisset · WWC 2023

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · WWC Europe 2026

1:39 min

Managing agentic infrastructure and non-deterministic hardware scheduling

Alejandro Saucedo Alejandro Saucedo · WWC 2025

3:23 min

The AI workload technology stack and its components

Lerna Ekmekcioglu Lerna Ekmekcioglu · Europe 2026 Virtual

Videos

See all

Related articles

See all