Support Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+20 more
Job description
- Serve as a senior technical escalation point, troubleshooting the hardest infrastructure and platform issues down to the hardware, driver, or kernel level when needed
- Quickly and accurately distinguish between hardware failures, driver issues, kernel-level problems, and customer workload misconfiguration, so issues get resolved correctly the first time
- Proactively identify process, tooling, and documentation gaps, and go fix them, not just wait for them to be assigned
- Use AI tools effectively to build scripts, automations, or small internal tools that close real operational gaps (no professional development background required)
- Perform root-cause analysis across distributed systems, clusters, and GPU infrastructure
- Craft clear documentation of solutions and contribute to evolving support procedures
- Collaborate closely with engineering teams to turn recurring customer pain points into permanent fixes
- Take escalations from peers while training and mentoring them in the process
- Participate in a rotating on-call schedule, owning major incidents and major customer issues
- Be ready to roll up your sleeves and pitch in wherever needed, especially during fast, high-volume deployments, Primary point-of-contact for a customer, providing on-site and remote HPC, Ethernet, and AI infrastructure support. Troubleshoot Linux systems, networking protocols, and interoperability; reproduce and resolve complex issues; collaborate with engineering, marketing, and support; document support methodologies and improve processes.
Requirements
- 3+ years of hands-on HPC experience in an administration, support, or engineering role.
- Very strong understanding and experience supporting Linux in a system administration role.
- Proven experience in HPC environments, showcasing your expertise in Linux cluster administration, with strong preference for Kubernetes and/or Slurm for cluster orchestration.
- Strong coding ability and CI/CD experience, with a track record of using AI-assisted tools to move fast.
- Proficiency with monitoring/logging tools (Prometheus, Grafana, Datadog).
- Strong skills in log analysis, debugging kernel-level issues, and performance profiling.
- Experience with CUDA, NCCL, NVLink, GPUDirect RDMA.
- Experience with high throughput networking technologies(IB/RoCE).
-
Knowledge of distributed AI/ML or HPC workloads.
-
Knowledge of TCP/IP, VPN, and firewalls in cloud environments.
- Ability to work independently and mentor junior support engineers.
Nice to Have
- Experience with virtualization and container (Docker, Kubernetes) technologies.
- Experience with neoclouds/GPU cloud providers.
- Flexible availability for potential shifts outside of normal working hours/weekends.
- Experience with high performance storage systems.
- Familiarity with infrastructure-as-code tools (Terraform, Ansible, etc.)
- Experience with Nvidia GPUs and Infiniband.
Benefits & conditions
Be an Early Applicant Remote Hiring Remotely in USA 122K-162K Annually Mid level Remote Hiring Remotely in USA 122K-162K Annually Mid level Provide senior-level escalation and troubleshooting for GPU/HPC infrastructure, diagnosing hardware, driver, and kernel issues. Perform root-cause analysis across clusters, build automations and docs, mentor peers, collaborate with engineering for permanent fixes, and participate in on-call rotation and high-volume deployments. The summary above was generated by AI
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda’s mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.
If you’d like to build the world’s best AI cloud, join us., This is a salaried exempt role. The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.
About Lambda
- Founded in 2012, with 500+ employees, and growing fast
- Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove
- We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG
- Our values are publicly available: https://lambda.ai/careers
- We offer generous cash & equity compensation
- Health, dental, and vision coverage for you and your dependents
- Wellness and commuter stipends for select roles
- 401k Plan with 2% company match (USA employees)
- Flexible paid time off plan that we all actually use, 122K-162K Annually Mid level 122K-162K Annually Mid level Software Provide senior-level technical escalation and support for GPU/HPC cloud infrastructure. Troubleshoot hardware, drivers, kernel, networking, and workload issues; perform root-cause analysis across clusters; build automations and documentation; mentor junior engineers; participate in on-call rotations and incident response; collaborate with engineering to implement permanent fixes. Top Skills: AnsibleCi/CdCudaDatadogDockerFirewallsGpudirect RdmaGrafanaInfinibandKubernetesLinuxNcclNvidia GpusNvlinkPrometheusRoceSlurmTcp/IpTerraformVpn NVIDIA, 108K-207K Annually Senior level 108K-207K Annually Senior level Artificial Intelligence * Computer Vision * Hardware * Robotics * Metaverse Provide onsite and remote technical support for NVIDIA Ethernet and AI infrastructure, troubleshoot Linux-based systems, debug networking and interoperability issues, collaborate with engineering and marketing, document support methodologies, and act as primary customer contact, spending at least one week per month onsite. Top Skills: AnsibleBashBgpChatgptCopilotCursorDockerEvpnGeminiGleanKubernetesLinuxNvidia Ethernet SwitchingNvidia Spectrum-XOspfPythonQosRoceTcpdumpVxlanWiresharkYaml
What you need to know about the Colorado Tech Scene
With a business-friendly climate and research universities like CU Boulder and Colorado State, Colorado has made a name for itself as a startup ecosystem. The state boasts a skilled workforce and high quality of life thanks to its affordable housing, vibrant cultural scene and unparalleled opportunities for outdoor recreation. Colorado is also home to the National Renewable Energy Laboratory, helping cement its status as a hub for renewable energy innovation.
Key Facts About Colorado Tech
- Number of Tech Workers: 260,000; 8.5% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Lockheed Martin, Century Link, Comcast, BAE Systems, Level 3
- Key Industries: Software, artificial intelligence, aerospace, e-commerce, fintech, healthtech
- Funding Landscape: $4.9 billion in VC funding in 2024 (Pitchbook)
- Notable Investors: Access Venture Partners, Ridgeline Ventures, Techstars, Blackhorn Ventures
- Research Centers and Universities: Colorado School of Mines, University of Colorado Boulder, University of Denver, Colorado State University, Mesa Laboratory, Space Science Institute, National Center for Atmospheric Research, National Renewable Energy Laboratory, Gottlieb Institute
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Highest Paying Tech Companies for Developers
7 Cloud Computing Trends Coming in 2025 for Developers
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence
Top 6 Hackathons for Developers in 2023