> Markdown version of [/jobs/ext/2716984-software-engineer-together-cloud-infrastructure](https://www.wearedevelopers.com/jobs/ext/2716984-software-engineer-together-cloud-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer - Together Cloud Infrastructure - **Company:** Together Ai - **Location:** San Francisco, CA, United States (Remote available) - **Experience:** Expert - **Salary:** $220,000.0 - $290,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Microsoft Azure, Cloud Computing, Nvidia CUDA, Data Centers, Memory Management, Fault Tolerance, Github, IBM Hardware Management Console, Hypervisor, Infrastructure as a Service (IaaS), InfiniBand, Virtual Private Networks (VPN), Open Source Technology, Overlay Transport Virtualization, Platform as a Service (PAAS), PCI Express, Ansible, Prometheus, Software Engineering, Virtual Local Area Networks, Graphics Processing Unit (GPU), Google Cloud, Grafana, Concurrency, Gpu Programming, Amazon Virtual Private Cloud (VPC), Build Management, AI Platforms, Kubernetes, Infrastructure Automation Frameworks, Terraform, Microservices - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/senior-software-engineer-together-cloud-infrastructure-together-ai-6826769 ## About the Role To be successful, you'll need to be deeply technical and possess excellent communication, collaboration, and diplomacy skills. You have strong fundamental software development skills. In addition, you have strong systems knowledge and troubleshooting abilities., * 5+ years of professional software development experience and proficiency in at least one backend programming language (Golang desired) * 5+ years experience writing high-performance, well-tested, production quality code, and demonstrated ownership of large-scale projects driven end-to-end to completion * Demonstrated experience with building and operating high-performance and/or globally distributed micro-service architectures across one or more cloud providers (AWS, Azure, GCP) * Excellent communication skills - able to write clear design docs and work effectively with both technical and non-technical team members * Strong systems knowledge across compute, networking, and storage, including concurrency, memory management, performant I/O, and scale * Experience building and operating reliable, customer-facing production systems at scale with infrastructure automation tools (Terraform, Ansible), monitoring/observability stacks (Prometheus, Grafana), and CI/CD pipelines (GitHub Actions, ArgoC, * Deep experience with Kubernetes internals, such as implementing non-trivial Kubernetes operators, device/storage/network plugins, custom schedulers, or patches to Kubernetes itself * Deep experience with VMs/hypervisors, such as QEMU/KVM, cloud-hypervisor, VFIO, virtio, PCIE passthrough, Kubevirt, SR-IOV * Deep experience with DC networking tech + solutions, such as VLAN, VXLAN, VPN, VPC, OVS/OVN * Experience with Cluster API or similar * Experience working on high-performance compute, networking, and/or storage * Experience virtualizing GPUs and/or Infiniband * Experience building IaaS or PaaS systems at scale * Experience with DPUs/SmartNICs * GPU programming, NCCL, CUDA knowledge ## Description * Design, build, and maintain performant, secure, and highly-available backend services/operators that run in our data centers and automate hardware management, such as Infiniband partitioning, in. DC parallel storage provisioning, and VM provisioning. * Design and build out the IaaS software layer for a new GB200 data center with thousands of GPUs. * Design and build distributed GPU scheduling and the global management plane that power on-demand and managed clusters across dozens of data centers * Develop infrastructure that powers our internal inference, RL, and fine-tuning products in addition to external cloud customers * Design and build systems that scale per-cluster capacity limits and automate the onboarding of new capacity * Work on a global multi-exabyte high-performance object store, serving massive datasets for pretraining and model weights for large-scale inference. * Build advanced observability stacks for our customers with automated node lifecycle management for fault-tolerant distributed pretraining and large-scale inference. * Perform architecture and research work for decentralized AI workloads * Work on the core, open-source Together AI platform * Create services, tools, and developer documentation * Create testing frameworks for robustness and fault-tolerance ## Related Videos - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Dev & Test in the Cloud? Deploy your cloud environments with Ansible & Terraform](https://www.wearedevelopers.com/videos/1607-dev-test-in-the-cloud-deploy-your-cloud-environments-with-ansible-terraform) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) - [Eclipse Che for Infrastructure Automation](https://www.wearedevelopers.com/videos/1611-eclipse-che-for-infrastructure-automation) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Why Attend a Developer Event in 2026?](https://www.wearedevelopers.com/magazine/688-why-attend-a-developer-event-in-2026)