Senior Platform Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+10 more
Job description
- Design and operate container orchestration platforms optimized for NVIDIA DGX/HGX-class hardware.
- Build bare-metal provisioning systems (PXE, Ironic, MAAS) to bring GPU clusters online at scale.
- Manage GPU lifecycle: driver stacks, CUDA/kernel compatibility, MIG slicing, and performance tuning.
- Partner with Network Engineering and DCOps to align physical infrastructure with software orchestration.
- Build automation and internal tooling in Go or Python to streamline cluster operations.
- Implement Terraform/Ansible-based IaC for fully auditable, repeatable infrastructure.
- Design high-resolution observability stacks (Prometheus/Grafana, DCGM, VictoriaMetrics).
- Participate in a specialized on-call rotation supporting GPU workloads and core platform services.
Requirements
Do you have experience in Systems engineering?, Requirements: 3+ years in Systems Engineering or HPC Infrastructure, strong Linux and bare-metal GPU experience, NVIDIA DGX/HGX, InfiniBand/RoCE, and automation with Python or Go, * 7+ years in systems, platform, or distributed systems engineering (10+ for Staff).
- Expert-level Linux knowledge: kernel modules, sysctl tuning, hugepages, container runtimes.
- Hands-on experience bootstrapping Kubernetes or SLURM on physical hardware.
- Strong proficiency in Go (preferred) or Python for systems-level automation.
- Deep familiarity with NVIDIA GPU ecosystems (drivers, CUDA, MIG).
- Working knowledge of InfiniBand or RoCEv2 networking and NCCL performance tuning.
- Experience building observability pipelines for hardware-accelerated environments.
- Ability to troubleshoot complex, multi-layered issues across hardware, networking, and orchestration.
- Strong cross-team communication - you’re the “glue” between Network, DCOps, and Software.
Bonus Points
- Experience with SLURM, Kubeflow, or distributed PyTorch.
- Integrating vendor APIs (NetBox, Vault, GitLab CI, etc.) into unified workflows.
- Infrastructure testing, chaos engineering, or cluster-level integration test suites.
- Designing telemetry aggregation across hardware, networking, and environmental systems.
Benefits & conditions
Pulled from the full job description
- Paid time off
- RSU, * $175k - $275k/year DOE
- RSU’s
- 5 weeks PTO
- 401k w/ match
- Comprehensive Benefit Plan
About the company
We build the high-performance, bare-metal GPU infrastructure that powers modern AI. Our team designs and operates large-scale NVIDIA DGX/HGX clusters, high-speed networking, and the automation that turns complex hardware into a reliable, production-ready platform. We work directly with the metal: provisioning nodes, tuning Linux, integrating InfiniBand/RoCE, and building the tooling that enables fast, secure, and scalable AI workloads.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on indeed.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Highest Paying Tech Companies for Developers
Top 6 Hackathons for Developers in 2023
Dev Digest 120 - Apple and peers
7 Cloud Computing Trends Coming in 2025 for Developers