> Markdown version of [/jobs/ext/1348712-technical-support-engineer-sysops](https://www.wearedevelopers.com/jobs/ext/1348712-technical-support-engineer-sysops). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Technical Support Engineer (SysOps) - **Company:** VAST, INC. - **Location:** Los Angeles, CA, United States - **Salary:** $90,000.0 - $130,000.0 - **Contract:** Permanent contract - **Skills:** Proxmox, Artificial Intelligence, Bash Shell, BIOS, Ubuntu (Operating System), CentOS, Command-Line Interface, Nvidia CUDA, System Configuration, Dynamic Host Configuration Protocol, Debian Linux, Software Debugging, Linux, Network Address Translation, Distributed Systems, Domain Name System (DNS), Firmware, General-Purpose Computing on Graphics Processing Units, Monitoring of Systems, Hypervisor, Image Management, Virtual Private Networks (VPN), Python (Programming Language), Kernel-Based Virtual Machine, Network Configuration and Change Management, Networking Basics, Red Hat Enterprise Linux, Tensorflow, Prometheus, Virtual Local Area Networks, Virtualization Technology, AI Infrastructure, Scripting, Pytorch, Grafana, Comptia Linux+, Firewall Services Module, Docker, Vmware - **Published:** July 19, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=b12452cf28bf1923 ## About the Role * Fluent in Linux - you navigate systems, read logs, and solve problems from the command line without hesitation * Methodical and thorough: you gather data, dig into root causes, and don't settle for surface-level fixes * A self-starter who can manage a queue of complex tickets with minimal supervision * Adaptable to a defined on-call rotation which may include weekend coverage * A clear written communicator: able to explain technical findings and write useful internal documentation * Genuinely curious about AI infrastructure, GPU computing, and distributed systems Must-Haves * Solid Linux SysOps experience: Ubuntu Server, RHEL/CentOS, Debian; comfortable with systems, networking, storage, and permissions * Proficiency with Docker: container debugging, Docker Compose, image management, cgroup resource limits, Docker storage/filesystem management * Experience with virtualization: Proxmox VE, VMware, or similar hypervisors; provisioning and troubleshooting VMs * Networking fundamentals: VLAN, DNS, DHCP, NAT, VPN, firewall rules, and general L2/L3 troubleshooting * Hands-on experience with NVIDIA GPU drivers, CUDA, and GPU workload troubleshooting (essential) * Scripting in Python and Bash for automation and diagnostic tooling * Strong English written communication: clear, professional, and technically precise * Experience providing technical support in a customer-facing or internal helpdesk context * Ability to prioritize across a concurrent queue of escalated tickets, triaging by severity and customer impact, balancing reactive resolution against proactive documentation and tooling work, and making clear judgment calls on when to escalate versus own resolution end-to-end Nice-to-Haves * Familiarity with AI/ML frameworks (TensorFlow, PyTorch) and running GPU-accelerated containers * Monitoring and observability experience (Prometheus, Grafana) * Relevant certifications: RHCSA, CompTIA Linux+, or similar * Knowledge of the Vast.ai platform as a client or infrastructure supplier ## Description This is a technical support role focused on escalated infrastructure issues that go beyond frontline triage. You'll be the engineering resource our L1 support team leans on when tickets get complex: diagnosing and resolving issues across the full stack - hardware/BIOS/firmware, networking, Ubuntu, Docker, NVIDIA CUDA/GPU, and virtualization (KVM). You'll handle higher-complexity issues, own escalation resolution end-to-end, and contribute to internal documentation and runbooks. The best engineers in this role don't just resolve tickets - they build the tooling and runbooks that eliminate recurring ones. You'll collaborate directly with the engineering team and host support team on systemic issues. Strong technical depth and support experience are the primary requirements. You should be comfortable working autonomously across Ubuntu environments, diagnosing container and GPU issues, and communicating findings clearly to both technical and non-technical audiences. Vast.ai users or hosts strongly preferred. This role is full-time and onsite in our office in Westwood (LA) Schedule: Sunday - Thursday., * Handle escalated support tickets, including GPU workload failures, container issues, networking problems, account infrastructure, and host-side configuration * Diagnose and resolve issues across Docker, NVIDIA CUDA/GPU drivers, and virtualization environments (KVM) * Troubleshoot network-layer issues: VLAN, DNS, DHCP, VPN, NAT, firewall rules, and connectivity failures on host machines * Investigate performance issues on GPU utilization, container resource constraints, thermal throttling, driver conflicts, disk I/O bottlenecks * Advise suppliers (hosts) on installation best practices - hardware setup, driver configuration, BIOS/firmware settings, and network configuration for optimal performance * Provide managed support for supplier onboarding and ongoing machine management, acting as a technical resource through installation, configuration, and post-setup troubleshooting * Write and maintain internal runbooks, escalation guides, and knowledge base articles to reduce repeat escalations * Build diagnostic and automation tooling in Python and Bash to reduce manual triage overhead * Collaborate with the engineering team and infrastructure support team to flag and document systemic or recurring platform issues * Assist clients and infrastructure suppliers working with AI frameworks (TensorFlow, PyTorch) and GPU-accelerated workloads * Provide coverage for L1 support team overflow during peak periods or incidents, per a defined on-call rotation ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [WebAssembly: The Next Frontier of Cloud Computing](https://www.wearedevelopers.com/videos/972-webassembly-the-next-frontier-of-cloud-computing) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [10M Data Records Lost, Underwater Computing, and Psychedelic Fish - Matthias Geniar](https://www.wearedevelopers.com/videos/1908-10m-data-records-lost-underwater-computing-and-psychedelic-fish-matthias-geniar) - [Generating code with Angular schematics](https://www.wearedevelopers.com/videos/129-generating-code-with-angular-schematics) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [6 Emerging Technologies We’ll Learn About in 2025](https://www.wearedevelopers.com/magazine/381-6-emerging-technologies-we-ll-learn-about-in-2025)