Software Engineer (Back-end)

TensorWave Inc.
United States
1 day ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

Intelligent Platform Management Interface Computer Clusters System Configuration Continuous Integration Firmware Github E2e Testing Ansible Prometheus Runbook Grafana Kubernetes
+5 more
Bare Metal Slurm Restful APIs Terraform Webhooks

Job description

We’re looking for a Software Engineer (Back-end) to join our platform team during an exciting phase of growth. In this role, you’ll own the end-to-end automation of provisioning, configuring, and operating large-scale GPU clusters across bare metal, Kubernetes, and Slurm environments. This is a hands-on technical role focused on building the tooling and pipelines that bring hundreds of GPU nodes online reliably and repeatably-working closely with cross-functional partners to support business objectives while upholding our standards for excellence, collaboration, and impact., * Build and maintain fully automated pipelines for provisioning bare metal GPU clusters from zero to production

  • Automate Slurm and Kubernetes cluster lifecycle-bootstrapping, upgrades, node provisioning, and decommissioning at scale
  • Develop and maintain infrastructure for GPU node configuration, including drivers and firmware
  • Own cluster validation pipelines, automating health checks and GPU burn-in tests
  • Build day-2 operations automation, including node remediation, rolling upgrades, and automated drain/cordon workflows
  • Write and maintain runbooks and documentation to enable reliable, repeatable operations
  • Own the full observability stack for automation services, provisioning pipelines, and cluster health systems

Requirements

  • 5+ years in infrastructure engineering or platform engineering
  • 3+ years writing production Go
  • Deep understanding of Kubernetes internals, including:
  • Informers and work queues
  • Controller-runtime and client-go
  • CRDs, custom controllers, and operators
  • Admission webhooks
  • Experience building Kubernetes Operators
  • Experience building gRPC and REST APIs in Go at production scale
  • Familiarity with bare metal infrastructure concepts, including PXE, IPMI, and BMC
  • Strong testing discipline across unit, integration, and end-to-end tests
  • Proven ownership of observability stacks such as Prometheus, Grafana, OpenTelemetry, and Loki (or similar)

Preferred Qualifications

  • Knowledge of GPU workload infrastructure
  • Experience with RoCE networking automation
  • Experience with GitOps tools such as ArgoCD
  • Experience with CI/CD tools such as GitHub Actions and Argo Workflows
  • Experience with Ansible and Terraform

Benefits & conditions

What We Offer

  • Stock Options

  • 100% paid Medical, Dental, and Vision insurance for Employees

  • Company Health Savings Account Contributions

  • 100% paid Short Term and Long Term Disability Insurance for Employees

  • Life and Voluntary Supplemental Insurance Options

  • Other Insurance Options, such as Pet & Legal Insurance

  • Various Supplementary Health Benefits, such as discounted Virtual Healthcare Appointments and Serious Illness Support

  • Flexible Spending Account

  • 401(k)

  • Employee Assistance Program

  • Flexible PTO

  • Paid Holidays

  • Parental Leave

  • Other In-Office Perks

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:42 min

Automating Skupper deployments using Ansible

Alex Soto Alex Soto · World Congress 2024

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

3:19 min

Executing complex workflows using Ansible Automation Platform

Goetz Rieger Goetz Rieger · World Congress 2025

2:40 min

Using GitHub primitives for internal documentation and corporate operations

Kyle Daigle · Coffee With Developers

Videos

See all

Related articles

See all