> Markdown version of [/jobs/ext/2977531-infrastructure-management-and-provisioning-engineer](https://www.wearedevelopers.com/jobs/ext/2977531-infrastructure-management-and-provisioning-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Infrastructure Management And Provisioning Engineer - **Company:** Roche - **Location:** Madrid, Spain - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Systems Engineering, Automation of Tests, Ubuntu (Operating System), Command-Line Interface, Cloud Computing, Configuration Management, Computer Engineering, Data Transmissions, DevOps, Red Hat Enterprise Linux, Ansible, Software Vulnerability Management, AI Infrastructure, High Performance Computing, System Availability, Delivery Pipeline, Technical Debt, Gitlab, Gitlab-ci, Git Flow, Infrastructure Automation Frameworks, Information Technology, Bare Metal, Performance Monitor - **Published:** September 18, 2026 - **Apply:** https://www.buscojobs.com.es/infrastructure-management-and-provisioning-engineer-en-madrid-ID-372174626 ## About the Role Education / ExperienceBachelor's or an advanced degree in Computer Science, Computer Engineering, or a similar technical discipline 5+ years of experience in systems engineering, DevOps, or platform infrastructure roles, with a proven track record of managing enterprise Linux environments at scale Deep, practical knowledge of operating system internals for both RHEL and Ubuntu OS Technical & Business SkillsAutomation & Orchestration: Advanced capability with Ansible on the command line and experience building scalable infrastructure pipelines using GitLab CI/CD Provisioning Tooling: Experience using NVIDIA Base Command Manager (Bright Cluster Manager) and Red Hat Image Builder (or related tools like Kickstart/Satellite) Modern Engineering Mindset: Strong adherence to git-based workflows, code-review methodologies, and infrastructure-as-code principles Troubleshooting Depth: Ability to isolate complex, multi-layered faults bridging hardware, kernel configurations, and automation scripts Leadership & MindsetLean & Agile Mindset: Passionate about continuous improvement, eliminating technical debt, and automating repetitive tasks to achieve scale Collaboration & Communication: Strong collaborative skills with an enterprise mindset, capable of working fluidly across team boundaries to drive platform success Intellectual Curiosity: Highly self-motivated to explore and adopt emerging technologies in the fast-evolving landscape of HPC and AI infrastructure engineering ## Description At Roche you can show up as yourself, embraced for the unique qualities you bring.Our culture encourages personal expression, open dialogue, and genuine connections, where you are valued, accepted and respected for who you are, allowing you to thrive both personally and professionally.This is how we aim to prevent, stop and cure diseases and ensure everyone has access to healthcare today and for generations to come.Join Roche, where every voice matters.The PositionJob DescriptionAs an Infrastructure Provisioning and Management Engineer within the Accelerated Compute Engineering (ACE) team, you will be responsible for overseeing and advancing our core infrastructure management and provisioning tech stack.This role has a strong focus on driving configuration-as-code, infrastructure-as-code (IaC), and modern automated provisioning best practices across our high-performance compute (HPC) and industry-leading AI Factory.You will own the lifecycle, deployment, and optimization of bare-metal and virtualized compute environments that power Roche's advanced computing initiatives.By treating infrastructure strictly as code and eliminating manual configurations, you will ensure our advanced clusters are highly reproducible, securely patched, and rapidly scalable to meet the evolving demands of computational science and large-scale AI workloads.Description Of The AreaHosting and Infrastructure (HI) provides mission?critical on?premise infrastructure, cloud hosting, connectivity, and technology products that enable all functions at every Roche site to develop, innovate, connect, and deliver compliant digital products across the Roche Enterprise.The Value Streams - Accelerated Compute Engineering (ACE) Team is focused on driving both customer success and platform success by acting as a center of excellence and delivery for the High Performance Compute and AI Infrastructure supporting AI and HPC use cases across Roche.This team facilitates seamless onboarding and adoption for business vertical customers needing accelerated compute-helping those infrastructure consumers with needs optimized for high availability, seamless data transfer, flexibility, speed, and the rapidly changing needs of AI-helping achieve rapid time-to-value.Job ResponsibilitiesAutomated Provisioning & Cluster OrchestrationDesign, deploy, and manage large-scale automated provisioning systems for multi-node HPC and AI Factory environmentsOwn and maintain the infrastructure management and provisioning tech stack underpinning the orchestration, monitoring, and provisioning of complex GPU and CPU workloadsStreamline bare-metal provisioning and node imaging pipelines to ensure minimal downtime and rapid expansion capabilitiesInfrastructure-as-Code (IaC) & Configuration GovernanceEnforce a strict configuration-as-code and infrastructure-as-code mindset, replacing manual interventions with repeatable automation scriptsAuthor, review, and maintain complex Ansible playbooks and roles for configuration management, patch deployment, and compliance drift remediationEstablish robust CI/CD pipelines using GitLab to test, validate, and deploy infrastructure changes safely across development, staging, and production clustersOperating System Engineering & Lifecycle ManagementIn partnership with Enterprise OS teams, standardize and manage operating system builds, with dual proficiency across HPC and AI Factory platformsUtilize solutions such as Red Hat Image Builder and NVIDIA Base Command Manager to create optimized, compliant, and secure custom golden images tailored for AI and high-performance computing workloadsManage OS lifecycles, including kernel tuning, automated package updates, and vulnerability management, ensuring alignment with global security standardsPlatform Reliability & CollaborationImplement proactive monitoring and alerting for infrastructure provisioning health, node availability, and configuration driftsAddress and help resolve complex, systemic infrastructure failures, contributing to post-mortem analyses to continuously improve platform resilienceQualificationsEducation / ExperienceBachelor's or an advanced degree in Computer Science, Computer Engineering, or a similar technical discipline5+ years of experience in systems engineering, DevOps, or platform infrastructure roles, with a proven track record of managing enterprise Linux environments at scaleDeep, practical knowledge of operating system internals for both RHEL and Ubuntu OSTechnical & Business SkillsAutomation & Orchestration: Advanced capability with Ansible on the command line and experience building scalable infrastructure pipelines using GitLab CI/CDProvisioning Tooling: Experience using NVIDIA Base Command Manager (Bright Cluster Manager) and Red Hat Image Builder (or related tools like Kickstart/Satellite)Modern Engineering Mindset: Strong adherence to git-based workflows, code-review methodologies, and infrastructure-as-code principlesTroubleshooting Depth: Ability to isolate complex, multi-layered faults bridging hardware, kernel configurations, and automation scriptsLeadership & MindsetLean & Agile Mindset: Passionate about continuous improvement, eliminating technical debt, and automating repetitive tasks to achieve scaleCollaboration & Communication: Strong collaborative skills with an enterprise mindset, capable of working fluidly across team boundaries to drive platform successIntellectual Curiosity: Highly self-motivated to explore and adopt emerging technologies in the fast-evolving landscape of HPC and AI infrastructure engineeringWho we areA healthier future drives us to innovate.Together, more than 100'000 employees across the globe are dedicated to advance science, ensuring everyone has access to healthcare today and for generations to come.Our efforts result in more than 26 million people treated with our medicines and over 30 billion tests conducted using our Diagnostics products.We empower each other to explore new possibilities, foster creativity, and keep our ambitions high, so we can deliver life-changing healthcare solutions that make a global impact.Let's build a healthier future, together.Roche is an Equal Opportunity Employer.#J-*****-Ljbffr ## Related Videos - [WeAreDevelopers LIVE - Modern DevOps for IoT Devices and More](https://www.wearedevelopers.com/videos/1805-wearedevelopers-live-modern-devops-for-iot-devices-and-more) - [Dev & Test in the Cloud? Deploy your cloud environments with Ansible & Terraform](https://www.wearedevelopers.com/videos/1607-dev-test-in-the-cloud-deploy-your-cloud-environments-with-ansible-terraform) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Enabling automated 1-click customer deployments with built-in quality and security](https://www.wearedevelopers.com/videos/83-enabling-automated-1-click-customer-deployments-with-built-in-quality-and-security) - [Eclipse Che for Infrastructure Automation](https://www.wearedevelopers.com/videos/1611-eclipse-che-for-infrastructure-automation) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers)