Software Engineer (Back-end)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+5 more
Job description
We’re looking for a Software Engineer (Back-end) to join our platform team during an exciting phase of growth. In this role, you’ll own the end-to-end automation of provisioning, configuring, and operating large-scale GPU clusters across bare metal, Kubernetes, and Slurm environments. This is a hands-on technical role focused on building the tooling and pipelines that bring hundreds of GPU nodes online reliably and repeatably-working closely with cross-functional partners to support business objectives while upholding our standards for excellence, collaboration, and impact., * Build and maintain fully automated pipelines for provisioning bare metal GPU clusters from zero to production
- Automate Slurm and Kubernetes cluster lifecycle-bootstrapping, upgrades, node provisioning, and decommissioning at scale
- Develop and maintain infrastructure for GPU node configuration, including drivers and firmware
- Own cluster validation pipelines, automating health checks and GPU burn-in tests
- Build day-2 operations automation, including node remediation, rolling upgrades, and automated drain/cordon workflows
- Write and maintain runbooks and documentation to enable reliable, repeatable operations
- Own the full observability stack for automation services, provisioning pipelines, and cluster health systems
Requirements
- 5+ years in infrastructure engineering or platform engineering
- 3+ years writing production Go
- Deep understanding of Kubernetes internals, including:
- Informers and work queues
- Controller-runtime and client-go
- CRDs, custom controllers, and operators
- Admission webhooks
- Experience building Kubernetes Operators
- Experience building gRPC and REST APIs in Go at production scale
- Familiarity with bare metal infrastructure concepts, including PXE, IPMI, and BMC
- Strong testing discipline across unit, integration, and end-to-end tests
- Proven ownership of observability stacks such as Prometheus, Grafana, OpenTelemetry, and Loki (or similar)
Preferred Qualifications
- Knowledge of GPU workload infrastructure
- Experience with RoCE networking automation
- Experience with GitOps tools such as ArgoCD
- Experience with CI/CD tools such as GitHub Actions and Argo Workflows
- Experience with Ansible and Terraform
Benefits & conditions
What We Offer
-
Stock Options
-
100% paid Medical, Dental, and Vision insurance for Employees
-
Company Health Savings Account Contributions
-
100% paid Short Term and Long Term Disability Insurance for Employees
-
Life and Voluntary Supplemental Insurance Options
-
Other Insurance Options, such as Pet & Legal Insurance
-
Various Supplementary Health Benefits, such as discounted Virtual Healthcare Appointments and Serious Illness Support
-
Flexible Spending Account
-
401(k)
-
Employee Assistance Program
-
Flexible PTO
-
Paid Holidays
-
Parental Leave
-
Other In-Office Perks
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Highest Paying Tech Companies for Developers
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence
Dev Digest 120 - Apple and peers
The Best X (Twitter) Accounts for Developers