> Markdown version of [/jobs/ext/1940578-ai-hpc-workload-orchestration-engineer](https://www.wearedevelopers.com/jobs/ext/1940578-ai-hpc-workload-orchestration-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Ai & Hpc Workload Orchestration Engineer - **Company:** Roche - **Location:** Madrid, Spain - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Systems Engineering, Distributed Systems, InfiniBand, Node.Js, Red Hat Enterprise Linux, System Availability, Multi-Agent Systems, Containerization, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Slurm - **Published:** August 5, 2026 - **Apply:** https://www.buscojobs.com.es/ai-hpc-workload-orchestration-engineer-en-madrid-ID-365536186 ## About the Role Education / Experience : Bachelor's or advanced degree in Computer Science, Applied Mathematics, Computational Engineering, or a similar technical discipline. 5+ years of systems engineering experience, with a heavy emphasis on workload scheduling, resource management, and cluster optimization for multi-tenant environments. Technical Proficiency : Deep familiarity with Enterprise Linux operating systems and distributed systems architecture. Expert-level proficiency in administering SLURM, including complex partition designs, accounting, and plug-in management. Highly proficient with Singularity for container runtime execution. AI Orchestration : Hands?on experience or deep architectural understanding of Run:ai, Kubernetes, and containerized GPU scheduling paradigms. Infrastructure Literacy : Solid understanding of high?speed interconnects (InfiniBand, RoCE) and multi?node communication architectures (MPI, NCCL) as they relate to job placement. Automation : Proficiency in automating scheduler configurations and telemetry gathering, or infrastructure automation tooling. Leadership & Mindset : Lean & Agile mindset; highly focused on driving efficiency, reducing idle compute time, and creating frictionless pathways for user workload submissions. Collaboration & Advocacy: Outstanding capability to translate scientific and AI model workflow challenges into scalable scheduler configurations. Intellectual Curiosity: A strong passion for remaining ahead of industry trends regarding GPU slicing, fractionalization, and the convergence of AI workloads with traditional HPC schedulers. ## Description At Roche you can show up as yourself, embraced for the unique qualities you bring.Our culture encourages personal expression, open dialogue, and genuine connections, where you are valued, accepted and respected for who you are, allowing you to thrive both personally and professionally.This is how we aim to prevent, stop and cure diseases and ensure everyone has access to healthcare today and for generations to come.Join Roche, where every voice matters.Job ResponsibilitiesOrchestration Stack Deployment & GovernanceDesign, implement, and maintain the SLURM Workload Manager ecosystem across HPC cluster architectures, ensuring high availability and optimal resource distribution.Deploy and manage Run:ai as the core orchestration and virtualization layer for the AI Factory, enabling fractional GPU allocation and dynamic resource allocation.Evaluate, architect, and implement SLURM Slinky integrations where required to seamlessly bridge Kubernetes-based AI orchestration with traditional HPC cluster resources.Containerization & Workload OptimizationDefine best practices and frameworks for containerized scientific execution, utilizing Singularity/Apptainer and/or Enroot to provide secure, reproducible performance environments for HPC.Translate user and workload requirements into optimized scheduling parameters (e.g., topology-aware scheduling, multi-node scaling).Actively profile and tune scheduling queues, QoS parameters, and fair-share policies to maximize multi-tenant efficiency.Platform Reliability & TelemetryPartner with Observability Engineers to implement continuous monitoring, telemetry, and reporting dashboards to track scheduler efficiency, queue wait times, and hardware utilization rates.Troubleshoot complex workload failures, including distributed training synchronization issues, MPI communication bottlenecks, and driver incompatibilities.Maintain configuration-as-code models for the scheduling tier, leveraging automation to deploy cluster policies uniformly.QualificationsEducation / Experience : Bachelor's or advanced degree in Computer Science, Applied Mathematics, Computational Engineering, or a similar technical discipline.5+ years of systems engineering experience, with a heavy emphasis on workload scheduling, resource management, and cluster optimization for multi-tenant environments.Technical Proficiency : Deep familiarity with Enterprise Linux operating systems and distributed systems architecture.Expert-level proficiency in administering SLURM, including complex partition designs, accounting, and plug-in management.Highly proficient with Singularity for container runtime execution.AI Orchestration : Hands?on experience or deep architectural understanding of Run:ai, Kubernetes, and containerized GPU scheduling paradigms.Infrastructure Literacy : Solid understanding of high?speed interconnects (InfiniBand, RoCE) and multi?node communication architectures (MPI, NCCL) as they relate to job placement.Automation : Proficiency in automating scheduler configurations and telemetry gathering, or infrastructure automation tooling.Leadership & Mindset : Lean & Agile mindset; highly focused on driving efficiency, reducing idle compute time, and creating frictionless pathways for user workload submissions.Collaboration & Advocacy: Outstanding capability to translate scientific and AI model workflow challenges into scalable scheduler configurations.Intellectual Curiosity: A strong passion for remaining ahead of industry trends regarding GPU slicing, fractionalization, and the convergence of AI workloads with traditional HPC schedulers.Roche is an Equal Opportunity Employer.#J-*****-Ljbffr ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [Stop using Node.js like in 2020! What changed and what you can do today with Node.js](https://www.wearedevelopers.com/videos/100011-stop-using-node-js-like-in-2020-what-changed-and-what-you-can-do-today-with-node-js) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Stop Using Node.js Like It’s 2020! - Alfonso Graziano](https://www.wearedevelopers.com/videos/1863-stop-using-node-js-like-it-s-2020-alfonso-graziano) - [Building AI Applications with LangChain and Node.js](https://www.wearedevelopers.com/videos/1512-building-ai-applications-with-langchain-and-node-js) ## Related Articles - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Top 6 Hackathons for Developers in 2023](https://www.wearedevelopers.com/magazine/263-top-6-hackathons-for-developers-in-2023) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)