Senior Platform Engineer

Everforth CyberCoders
San Francisco, CA, United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Compensation
$175,000.0 - $275,000.0
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Systems Engineering Build Automation Computer Clusters Nvidia CUDA Linux Distributed Systems InfiniBand Python (Programming Language) Performance Tuning Reliability Engineering Ansible
+10 more
Prometheus Systems Integration Pytorch Grafana Gitlab-ci Integration Tests Kubernetes Bare Metal Slurm Terraform

Job description

  • Design and operate container orchestration platforms optimized for NVIDIA DGX/HGX-class hardware.
  • Build bare-metal provisioning systems (PXE, Ironic, MAAS) to bring GPU clusters online at scale.
  • Manage GPU lifecycle: driver stacks, CUDA/kernel compatibility, MIG slicing, and performance tuning.
  • Partner with Network Engineering and DCOps to align physical infrastructure with software orchestration.
  • Build automation and internal tooling in Go or Python to streamline cluster operations.
  • Implement Terraform/Ansible-based IaC for fully auditable, repeatable infrastructure.
  • Design high-resolution observability stacks (Prometheus/Grafana, DCGM, VictoriaMetrics).
  • Participate in a specialized on-call rotation supporting GPU workloads and core platform services.

Requirements

Do you have experience in Systems engineering?, Requirements: 3+ years in Systems Engineering or HPC Infrastructure, strong Linux and bare-metal GPU experience, NVIDIA DGX/HGX, InfiniBand/RoCE, and automation with Python or Go, * 7+ years in systems, platform, or distributed systems engineering (10+ for Staff).

  • Expert-level Linux knowledge: kernel modules, sysctl tuning, hugepages, container runtimes.
  • Hands-on experience bootstrapping Kubernetes or SLURM on physical hardware.
  • Strong proficiency in Go (preferred) or Python for systems-level automation.
  • Deep familiarity with NVIDIA GPU ecosystems (drivers, CUDA, MIG).
  • Working knowledge of InfiniBand or RoCEv2 networking and NCCL performance tuning.
  • Experience building observability pipelines for hardware-accelerated environments.
  • Ability to troubleshoot complex, multi-layered issues across hardware, networking, and orchestration.
  • Strong cross-team communication - you’re the “glue” between Network, DCOps, and Software.

Bonus Points

  • Experience with SLURM, Kubeflow, or distributed PyTorch.
  • Integrating vendor APIs (NetBox, Vault, GitLab CI, etc.) into unified workflows.
  • Infrastructure testing, chaos engineering, or cluster-level integration test suites.
  • Designing telemetry aggregation across hardware, networking, and environmental systems.

Benefits & conditions

Pulled from the full job description

  • Paid time off
  • RSU, * $175k - $275k/year DOE
  • RSU’s
  • 5 weeks PTO
  • 401k w/ match
  • Comprehensive Benefit Plan

About the company

We build the high-performance, bare-metal GPU infrastructure that powers modern AI. Our team designs and operates large-scale NVIDIA DGX/HGX clusters, high-speed networking, and the automation that turns complex hardware into a reliable, production-ready platform. We work directly with the metal: provisioning nodes, tuning Linux, integrating InfiniBand/RoCE, and building the tooling that enables fast, secure, and scalable AI workloads.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · WWC Europe 2026

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · WWC Europe 2026

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · WWC 2025

2:39 min

Experiencing core Linux capabilities for DevOps administration

Michael Cade · LIVE

Videos

See all

Related articles

See all