> Markdown version of [/jobs/ext/2729064-infrastructure-engineer-platform](https://www.wearedevelopers.com/jobs/ext/2729064-infrastructure-engineer-platform). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Infrastructure Engineer - Platform - **Company:** Nvidia - **Location:** Berlin, Germany - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Intelligent Platform Management Interface, Bash Shell, Border Gateway Protocol, BIOS, Ubuntu (Operating System), Nvidia CUDA, Continuous Integration, Data Centers, Linux, DevOps, RAID, Firmware, InfiniBand, Data Intelligence, Python (Programming Language), Node.Js, Performance Tuning, Remote Direct Memory Access, Red Hat Enterprise Linux, Ansible, Ceph (Software), Network Switches, Scripting, Extensible Firmware Interface, Large Language Models, Git, Kubernetes, Infrastructure Automation Frameworks, Bare Metal, Machine Learning Operations, TensorRT, Terraform, Network Server, Nvme - **Published:** September 5, 2026 - **Apply:** https://startup.jobs/infrastructure-systems-engineer-orcrist-technologies-8276147 ## About the Role * 5+ years in bare-metal, data-center, or systems infrastructure engineering, with hands-on ownership of physical and compute infrastructure at scale. * Strong bare-metal Linux (Ubuntu, RKE2, Talos, or similar): provisioning, firmware/BMC management, PXE/iPXE, kernel and storage tuning, systemd, RAID. * Real experience with infrastructure automation (Ansible, Terraform, or equivalent), Git and CI/CD, and scripting in Python, Bash, or Go. * Solid, current data-center networking fundamentals: L2/L3, IP Fabric, BGP, and switching - this is a hard requirement, not a nice-to-have. RDMA (RoCE/InfiniBand) experience is a strong plus. * Comfortable operating in air-gapped or on-prem environments and traveling to customer sites for builds and deployments. * Practical hardware sizing literacy: power budgets, thermal/cooling, UPS sizing, and rack integration. * Documentation-focused, methodical, and calm during hardware incidents. Eligible to work in Germany., * German language (B1+); exposure to regulated or security-critical environments (e.g. BSI C5, ISO 27001, or defense-sector delivery) is a plus. * NVIDIA GPU stack knowledge (drivers, CUDA, GPU Operator, MIG, DCGM) and cross-node GPU interconnect experience (NVLink, InfiniBand, NCCL). * Kubernetes bare-metal fundamentals - how cluster bring-up and GPU device plugins interact with the underlying hardware, not day-to-day cluster operation. * Inference optimization (vLLM, TensorRT-LLM, quantization) and familiarity with switch NOS (SONiC/Cumulus). * Relevant certifications (NVIDIA, Red Hat, CKA/CKS) or field/forward-deployed engineering experience. ## Description You'll own the layer everything else runs on: bare-metal servers, operating systems, data-center networking, and storage across on-prem and fully air-gapped sites - the physical and infrastructure foundation our platform and GPU fleets are built on. You design, build, and operate server fleets with a strong automation and DevOps mindset, then partner with our SRE, MLOps, and ML teams to ensure everything running above the metal - including GPU inference - performs reliably at scale. Some of this work is hands-on at customer sites, where you size, rack, and commission self-contained server environments with no internet uplink. We weight depth in modern data-center infrastructure, networking, and automation more heavily than GPU-specific experience. A strong infrastructure and network engineer with a genuine automation mindset - even without prior GPU exposure - is a better fit for this role than a candidate with GPU expertise whose networking background is rooted in legacy, corporate-style L2 designs., * Design, size, provision, and operate bare-metal server fleets across on-prem and air-gapped environments (firmware/BIOS/UEFI, BMC via Redfish/IPMI, OS, RAID, kernel and storage tuning) using zero-touch provisioning (PXE/iPXE, MAAS/Metal3/Tinkerbell/Ironic) and automation (Ansible, Terraform, or equivalent - the tooling matters less than the automation mindset). * Build and run modern data-center networking: L2/L3 design, IP Fabric, BGP and switching, and RDMA fabrics (RoCE/InfiniBand) sized to scale without ripping out the core. * Engineer resilient, highly available storage (Ceph/Rook, NVMe) with capacity planning and encryption at rest. * Operate confidently in air-gapped and on-prem environments: offline mirrors and registries, signed artifacts, firmware/driver lifecycle without internet access, and system hardening. * Support MLOps and inference-serving fundamentals - GPU model serving (Triton/KServe/vLLM), GPU scheduling and sharing, and throughput/latency optimization - in partnership with our SRE and ML teams. * Plan and run on-site build-outs: rack integration, power budgets, thermal/cooling and UPS sizing, commissioning, capacity planning, runbooks, and operator handover, with SWaP awareness for field sites. ## Related Videos - [A Deep Dive on How To Leverage the NVIDIA GB200 for Ultra-Fast Training and Inference on Kubernetes](https://www.wearedevelopers.com/videos/1625-a-deep-dive-on-how-to-leverage-the-nvidia-gb200-for-ultra-fast-training-and-inference-on-kubernetes) - [10M Data Records Lost, Underwater Computing, and Psychedelic Fish - Matthias Geniar](https://www.wearedevelopers.com/videos/1908-10m-data-records-lost-underwater-computing-and-psychedelic-fish-matthias-geniar) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [How to Submit CFPs and Get into Public Speaking - Moran Weber](https://www.wearedevelopers.com/videos/2112-how-to-submit-cfps-and-get-into-public-speaking-moran-weber) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Fullstack developer salary in Germany [2023]](https://www.wearedevelopers.com/magazine/197-fullstack-developer-salary-in-germany-2023)