> Markdown version of [/jobs/ext/2708735-software-engineer-back-end](https://www.wearedevelopers.com/jobs/ext/2708735-software-engineer-back-end). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer (Back-end) - **Company:** TensorWave Inc. - **Location:** United States (Remote available) - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Intelligent Platform Management Interface, Computer Clusters, System Configuration, Continuous Integration, Firmware, Github, E2e Testing, Ansible, Prometheus, Runbook, Grafana, Kubernetes, Bare Metal, Slurm, Restful APIs, Terraform, Webhooks - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/software-engineer-back-end-tensorwave-7810650 ## About the Role * 5+ years in infrastructure engineering or platform engineering * 3+ years writing production Go * Deep understanding of Kubernetes internals, including: + Informers and work queues + Controller-runtime and client-go + CRDs, custom controllers, and operators + Admission webhooks * Experience building Kubernetes Operators * Experience building gRPC and REST APIs in Go at production scale * Familiarity with bare metal infrastructure concepts, including PXE, IPMI, and BMC * Strong testing discipline across unit, integration, and end-to-end tests * Proven ownership of observability stacks such as Prometheus, Grafana, OpenTelemetry, and Loki (or similar) Preferred Qualifications * Knowledge of GPU workload infrastructure * Experience with RoCE networking automation * Experience with GitOps tools such as ArgoCD * Experience with CI/CD tools such as GitHub Actions and Argo Workflows * Experience with Ansible and Terraform ## Description We're looking for a Software Engineer (Back-end) to join our platform team during an exciting phase of growth. In this role, you'll own the end-to-end automation of provisioning, configuring, and operating large-scale GPU clusters across bare metal, Kubernetes, and Slurm environments. This is a hands-on technical role focused on building the tooling and pipelines that bring hundreds of GPU nodes online reliably and repeatably-working closely with cross-functional partners to support business objectives while upholding our standards for excellence, collaboration, and impact., * Build and maintain fully automated pipelines for provisioning bare metal GPU clusters from zero to production * Automate Slurm and Kubernetes cluster lifecycle-bootstrapping, upgrades, node provisioning, and decommissioning at scale * Develop and maintain infrastructure for GPU node configuration, including drivers and firmware * Own cluster validation pipelines, automating health checks and GPU burn-in tests * Build day-2 operations automation, including node remediation, rolling upgrades, and automated drain/cordon workflows * Write and maintain runbooks and documentation to enable reliable, repeatable operations * Own the full observability stack for automation services, provisioning pipelines, and cluster health systems ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Dev & Test in the Cloud? Deploy your cloud environments with Ansible & Terraform](https://www.wearedevelopers.com/videos/1607-dev-test-in-the-cloud-deploy-your-cloud-environments-with-ansible-terraform) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Eclipse Che for Infrastructure Automation](https://www.wearedevelopers.com/videos/1611-eclipse-che-for-infrastructure-automation) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [The Best X (Twitter) Accounts for Developers](https://www.wearedevelopers.com/magazine/294-the-best-x-twitter-accounts-for-developers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)