> Markdown version of [/jobs/ext/3077373-staff-software-engineer-ai-compute-together-cloud](https://www.wearedevelopers.com/jobs/ext/3077373-staff-software-engineer-ai-compute-together-cloud). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Software Engineer - AI Compute, Together Cloud - **Company:** Together Ai - **Location:** San Francisco, CA, United States (Remote available) - **Experience:** Expert - **Salary:** $260,000.0 - $300,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Computing Platforms, Microsoft Azure, Software as a Service, Cloud Computing, Nvidia CUDA, Continuous Integration, Data Centers, Software Design Documents, Distributed Systems, Memory Management, Expert Systems, Fault Tolerance, Github, Hypervisor, Infrastructure as a Service (IaaS), InfiniBand, Virtual Private Networks (VPN), Network Virtualization, Open Source Technology, Overlay Transport Virtualization, Platform as a Service (PAAS), PCI Express, Ansible, Prometheus, Software Engineering, Virtual Local Area Networks, Graphics Processing Unit (GPU), Google Cloud, Grafana, Concurrency, Gpu Programming, Amazon Virtual Private Cloud (VPC), Backend, AI Platforms, Kubernetes, Production Code, Slurm, Hardware Infrastructure, Terraform, Golang, Microservices - **Published:** September 25, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=f19a35680448679b ## About the Role * 7+ years of professional software development experience, with expert-level proficiency in at least one backend language (Golang desired), writing high-performance, well-tested, production-quality code. * Track record of owning the architecture of large distributed systems from blank page to production at scale, including the judgment calls that could not be reversed cheaply. * Deep experience building and operating globally distributed, high-performance microservice architectures across one or more cloud providers (AWS, Azure, GCP). * Expert systems knowledge across compute, networking, and storage - including concurrency, memory management, performant I/O, and scale at a global level. * Demonstrated technical leadership beyond your own commits: mentoring senior engineers, leading design reviews, and driving alignment across teams that do not report to you. * Excellent communication and diplomacy skills - able to write design docs that settle arguments, and to work effectively with technical and non-technical stakeholders. * Experience building and operating reliable, customer-facing production systems at scale, and owning the infrastructure automation (Terraform, Ansible), observability (Prometheus, Grafana), and CI/CD (GitHub Actions, ArgoCD) that keep them healthy., * Deep Kubernetes internals experience, such as implementing non-trivial Kubernetes operators, device/storage/network plugins, custom schedulers, or patches to Kubernetes itself * Deep experience with VMs/hypervisors, such as QEMU/KVM, cloud-hypervisor, VFIO, virtio, PCIE passthrough, Kubevirt, SR-IOV * Deep experience with DC networking tech + solutions, such as VLAN, VXLAN, VPN, VPC, OVS/OVN * Experience with Cluster API or similar * Experience working on high-performance compute, networking, and/or storage * Experience virtualizing GPUs and/or InfiniBand * Experience building IaaS or PaaS systems at scale * Experience with DPUs/SmartNICs * GPU programming, NCCL, CUDA knowledge ## Description As a Staff Software Engineer focusing on AI Compute in the Together Cloud org, you will set technical direction for and build major components of the next generation AI cloud platform - a highly available, global cloud infrastructure with cutting-edge virtualization of the latest ML hardware: GB300s/VRs, BlueField DPUs, InfiniBand and dual/quad-plane RoCEv2 fabrics. That virtualized computing platform powers our own SaaS products - inference, RL, and fine-tuning - and serves external cloud customers through self-serve offerings such as on-demand/reserved Kubernetes/Slurm clusters, across dozens of data centers and hundreds of thousands of GPUs. This is an architect-and-build role. Fully automated bootstrapping of GPU data centers, high-performance virtualization of GPU compute and DC networking without compromising isolation or portability, and fault-tolerant decentralized control planes - you'll set the architecture for these across our global and in-DC services, and be a key owner of the hardest parts, in the code as well as the design. Your designs will span the IaaS layer of a greenfield Vera Rubin data center up to the global management plane that schedules capacity across all of them. At this level the job is as much leverage as code: the standards you set and the engineers you grow decide how fast the rest of Together Cloud ships., * Own the GPU and network virtualization stack: the hypervisor, kernel, and SDN work that keeps GPU compute and DC networking high-performance, portable, and strongly isolated across heterogeneous hardware. * Own the in-DC IaaS layer: architect and roadmap the services, Kubernetes operators, and libraries that provision and manage compute, storage, and networks in our data centers - VMs, parallel filesystems, VPCs, and InfiniBand partitions; lead its build-out for a new Vera Rubin data center with thousands of GPUs, from hardware bring-up to customer-facing API. * Design the GPU scheduling and global management plane: the distributed control plane behind on-demand and reserved clusters across dozens of data centers, including the systems that scale per-cluster limits and automate the onboarding of new capacity. * Architect monitoring and automated remediation for fault tolerance: the strategy for automated detection, isolation, and recovery of failed nodes that keeps distributed pretraining and large-scale inference running through hardware failures. * Set technical direction across teams: lead design reviews, resolve cross-cutting architectural disagreements, unblock cross-team dependencies and integration risks, and define the standards other engineers build against - measured in cluster reliability, time-to-first-GPU on new capacity, and quality at scale. * Grow the team: mentor senior and junior engineers, deepen the team's expertise in virtualization, DC networking, and GPU infrastructure, and help raise the hiring bar for Together Cloud. * Set the engineering bar: create the testing frameworks, tools, and developer documentation that make our systems robust and usable by other teams, and shape the core, open-source Together AI platform. To be successful you'll need to be deeply technical and an excellent communicator - expert software development fundamentals, deep systems knowledge and troubleshooting instincts, and the leadership and diplomacy skills to align teams that don't report to you. Much of this work starts ambiguous, and we expect you to define the scope yourself and drive it to production. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) - [Scoring 2000 Products per Request: Performance Pitfalls in Golang](https://www.wearedevelopers.com/videos/2073-scoring-2000-products-per-request-performance-pitfalls-in-golang) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)