> Markdown version of [/jobs/ext/2617689-technical-support-engineer-inference-us-weekends](https://www.wearedevelopers.com/jobs/ext/2617689-technical-support-engineer-inference-us-weekends). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Technical Support Engineer (Inference) - US Weekends - **Company:** Together Ai - **Location:** San Francisco, CA, United States (Remote available) - **Salary:** $160,000.0 - $230,000.0 - **Contract:** Permanent contract - **Skills:** JavaScript (Programming Language), Artificial Intelligence, Amazon Web Services, Microsoft Azure, Cloud Computing, Computer Clusters, Complex Networks, Software Debugging, DevOps, Python (Programming Language), CURL, Performance Tuning, Ansible, Prometheus, TypeScript, Weka, Datadog, Data Storage Management, High Performance Computing, Postman, Large Language Models, Grafana, Generative AI, Git Flow, Kubernetes, Bug Reporting, Slurm, Restful APIs - **Published:** August 4, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=3e732302c5e6b246 ## About the Role * 6+ years of experience in a customer-facing technical role, SRE, DevOps, or infrastructure engineering, with at least 1 year in a support role for an AI service * Experience as an SRE or DevOps engineer working with Kubernetes * Strong technical background, with knowledge of AI, ML, GPU technologies and their integration into high-performance computing (HPC) environments. * Advanced, production-level experience with infrastructure services (e.g., Kubernetes, SLURM), infrastructure as code solutions (e.g., Ansible) high-performance network fabrics, NFS-based storage management, and container infrastructure * Familiarity with operating storage systems in HPC environments such as Vast and Weka * Proven ability to diagnose complex network-layer issues and read traces * Strong knowledge of Python, TypeScript, and/or JavaScript with testing/debugging experience using curl and Postman-like tools * Demonstrated expertise with observability tooling (e.g., Prometheus, Grafana) at scale * Deep familiarity with REST API debugging and HTTP semantics * Experience with LLM inference frameworks and LoRA fine-tuning and common training failure modes * Experience with Infrastructure as Code and Git-based workflows * Background in GPU cluster management * Cloud platform experience (AWS, GCP, and/or Azure) * Foundational understanding in the installation, configuration, administration, troubleshooting, and securing of compute clusters. * Complex technical problem solving and troubleshooting, with a proactive approach to issue resolution * Ability to work cross-functionally with teams such as Sales, Engineering, Support, Product and Research to drive customer success. * Strong sense of ownership and willingness to learn new skills to ensure both team and customer success. * Excellent communication and interpersonal skills, with the ability to explain complex technical concepts to non-technical stakeholders. * Ability to operate in dynamic environments, adept at managing multiple projects, and comfortable with frequent context switching and prioritization. ## Description As a Technical Support Engineer at a pioneering AI company, you'll be the first line of defense to support customers as they build out training, fine tuning, and inference solutions with Together AI. You'll dive deep into complex technical challenges, providing swift and effective solutions while serving as a product expert. As a part of the Customer Experience organization, you will collaborate closely with product and sales, driving continuous improvement of our offerings. This is an exciting opportunity for a deeply technical professional passionate about AI and customer success to make a significant impact in a fast-paced, innovative environment., * Engage directly with customers to tackle and resolve complex technical challenges involving our cutting-edge GPU clusters and our inference and fine-tuning services; ensure swift and effective solutions every time. * Act as a customer facing SRE to ensure our customer's Inference endpoints (running on Kubernetes) remain healthy, stable, and performant * Become a product expert in all of our Gen AI solutions, serving as the last line of technical defense before issues are escalated to Engineering and Product teams. * Assist with hardware and platform migrations by validating system health and traffic routing. Monitor dashboards to detect anomalies and escalate with data-backed analysis * Manage customer-facing communications during incidents and degradations; translate deep technical findings (latency regressions, provider issues, network reachability drops) into clear, evidence-backed updates without exposing platform internals * Contribute infrastructure changes for model deployment, capacity rebalancing, and cluster configuration. You will execute infrastructure changes via pull requests (infra-as-code) for tasks such as endpoint configuration, model bringup/bringdown, and capacity scaling * Flag engine-level bugs with logs and reproduction steps for engineering * Collaborate seamlessly across Engineering, Research, and Product teams to address customer concerns; collaborate with senior leaders both internally and externally to ensure the highest levels of customer satisfaction. * Transform customer insights into action by identifying patterns in support cases and working with Engineering and Go-To-Market teams to drive Together's roadmap (e.g., future models to support) * Maintain detailed documentation of system configurations, procedures, troubleshooting guides, and FAQs to facilitate knowledge sharing with team and customers. * Be flexible in providing support coverage during holidays, nights and weekends as required by business needs to ensure consistent and reliable service for our customers. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Don’t Insert Crazy! On cURL and AI Slop - Daniel Stenberg](https://www.wearedevelopers.com/videos/1796-don-t-insert-crazy-on-curl-and-ai-slop-daniel-stenberg) - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) - [Capture the Flag 101](https://www.wearedevelopers.com/videos/416-capture-the-flag-101) - [Coffee with Developers - Maria Apazoglou](https://www.wearedevelopers.com/videos/1209-coffee-with-developers-maria-apazoglou) ## Related Articles - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud)