> Markdown version of [/jobs/ext/3062747-software-engineer-ai-compute](https://www.wearedevelopers.com/jobs/ext/3062747-software-engineer-ai-compute). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer - AI Compute - **Company:** Together Ai - **Location:** San Francisco, CA, United States (Remote available) - **Experience:** Expert - **Salary:** $220,000.0 - $270,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Microsoft Azure, Cloud Computing, Code Review, Nvidia CUDA, Continuous Integration, Data Centers, Programming Tools, Memory Management, Fault Tolerance, Github, Hypervisor, Infrastructure as a Service (IaaS), InfiniBand, Virtual Private Networks (VPN), Open Source Technology, Overlay Transport Virtualization, Platform as a Service (PAAS), PCI Express, Cloud Services, Ansible, Prometheus, Software Engineering, Virtual Local Area Networks, Graphics Processing Unit (GPU), Google Cloud, Grafana, Concurrency, Gpu Programming, Amazon Virtual Private Cloud (VPC), AI Platforms, Kubernetes, Production Code, Terraform, Golang, Microservices - **Published:** September 25, 2026 - **Apply:** https://startup.jobs/senior-software-engineer-ai-compute-together-cloud-together-ai-6826769 ## About the Role To be successful, you'll need to be deeply technical, ready to own projects truly end-to-end, and an excellent communicator - strong software development fundamentals, strong systems knowledge and troubleshooting instincts, and the collaboration and diplomacy skills to work across teams., * 5+ years of professional software development experience, with strong proficiency in at least one backend programming language (Golang desired). * Demonstrated ownership of large-scale projects driven end-to-end to completion, writing high-performance, well-tested, production-quality code. * Demonstrated experience building and operating high-performance and/or globally distributed micro-service architectures across one or more cloud providers (AWS, Azure, GCP). * Strong systems knowledge across compute, networking, and storage, including concurrency, memory management, performant I/O, and scale. * Excellent communication skills - able to write clear design docs and work effectively with both technical and non-technical team members. * Experience building and operating reliable, customer-facing production systems at scale, and owning the infrastructure automation (Terraform, Ansible), observability (Prometheus, Grafana), and CI/CD (GitHub Actions, ArgoCD) that keep them healthy., * Proficiency with Kubernetes internals, such as implementing non-trivial Kubernetes operators, device/storage/network plugins, custom schedulers, or patches to Kubernetes itself * Proficiency with VMs/hypervisors, such as QEMU/KVM, cloud-hypervisor, VFIO, virtio, PCIE passthrough, Kubevirt, SR-IOV * Proficiency with DC networking tech + solutions, such as VLAN, VXLAN, VPN, VPC, OVS/OVN * Experience with Cluster API or similar * Experience working on high-performance compute, networking, and/or storage * Experience virtualizing GPUs and/or InfiniBand * Experience building IaaS or PaaS systems at scale * Experience with DPUs/SmartNICs * GPU programming, NCCL, CUDA knowledge ## Description * Build the GPU and network virtualization stack: the hypervisor, kernel, and SDN work that makes GPU compute and DC networking high-performance, portable, and strongly isolated across heterogeneous hardware. * Build and maintain our in-DC IaaS layer: the services, Kubernetes operators, and libraries that provision and manage compute, storage, and networks in our data centers, including VMs, parallel filesystems, VPCs, and InfiniBand partitions. Implement and harden the bring-up path for a new Vera Rubin data center with thousands of GPUs. * Scale the distributed GPU scheduling and global management plane: the control-plane services behind on-demand and reserved clusters, including the automation that onboards new capacity and raises per-cluster limits. * Harden the monitoring and automated remediation layer for fault tolerance: automated detection, isolation, and recovery of failed nodes that keeps distributed pretraining and large-scale inference fault-tolerant. * Own your components end-to-end: write the design docs, break the work into milestones that ship incrementally, and improve the reliability of what's already in production. * Raise the bar around you: code review, design feedback, and mentoring junior engineers. * Build the tooling other teams rely on: testing frameworks, developer tools, and documentation that make our systems robust and usable across teams, plus contributions to the core, open-source Together AI platform. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Building the Nervous System of AI - Michael Kagan (NVIDIA)](https://www.wearedevelopers.com/videos/2133-building-the-nervous-system-of-ai-michael-kagan-nvidia) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence)