Support Engineer

Lambda Inc.
United States
about 1 month ago
Apply on jobs.ashbyhq.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Compensation
$122,000.0 - $162,000.0
Working hours
Shift work

Tech stack

Artificial Intelligence Computer Clusters Profiling Nvidia CUDA Continuous Integration Software Debugging Linux Ethernet InfiniBand Interoperability Virtual Private Networks (VPN) Linux System Administration
+20 more
Log Analysis Machine Learning Remote Direct Memory Access Ansible Prometheus TCP/IP AI Infrastructure Datadog Data Logging Graphics Processing Unit (GPU) Cloud Platform System High Performance Computing Computer Network Technologies Grafana Firewalls (Computer Science) Kubernetes Infrastructure Automation Frameworks Slurm Terraform Docker

Job description

  • Serve as a senior technical escalation point, troubleshooting the hardest infrastructure and platform issues down to the hardware, driver, or kernel level when needed
  • Quickly and accurately distinguish between hardware failures, driver issues, kernel-level problems, and customer workload misconfiguration, so issues get resolved correctly the first time
  • Proactively identify process, tooling, and documentation gaps, and go fix them, not just wait for them to be assigned
  • Use AI tools effectively to build scripts, automations, or small internal tools that close real operational gaps (no professional development background required)
  • Perform root-cause analysis across distributed systems, clusters, and GPU infrastructure
  • Craft clear documentation of solutions and contribute to evolving support procedures
  • Collaborate closely with engineering teams to turn recurring customer pain points into permanent fixes
  • Take escalations from peers while training and mentoring them in the process
  • Participate in a rotating on-call schedule, owning major incidents and major customer issues
  • Be ready to roll up your sleeves and pitch in wherever needed, especially during fast, high-volume deployments, Primary point-of-contact for a customer, providing on-site and remote HPC, Ethernet, and AI infrastructure support. Troubleshoot Linux systems, networking protocols, and interoperability; reproduce and resolve complex issues; collaborate with engineering, marketing, and support; document support methodologies and improve processes.

Requirements

  • 3+ years of hands-on HPC experience in an administration, support, or engineering role.
  • Very strong understanding and experience supporting Linux in a system administration role.
  • Proven experience in HPC environments, showcasing your expertise in Linux cluster administration, with strong preference for Kubernetes and/or Slurm for cluster orchestration.
  • Strong coding ability and CI/CD experience, with a track record of using AI-assisted tools to move fast.
  • Proficiency with monitoring/logging tools (Prometheus, Grafana, Datadog).
  • Strong skills in log analysis, debugging kernel-level issues, and performance profiling.
  • Experience with CUDA, NCCL, NVLink, GPUDirect RDMA.
  • Experience with high throughput networking technologies(IB/RoCE).
  • Knowledge of distributed AI/ML or HPC workloads.

  • Knowledge of TCP/IP, VPN, and firewalls in cloud environments.

  • Ability to work independently and mentor junior support engineers.

Nice to Have

  • Experience with virtualization and container (Docker, Kubernetes) technologies.
  • Experience with neoclouds/GPU cloud providers.
  • Flexible availability for potential shifts outside of normal working hours/weekends.
  • Experience with high performance storage systems.
  • Familiarity with infrastructure-as-code tools (Terraform, Ansible, etc.)
  • Experience with Nvidia GPUs and Infiniband.

Benefits & conditions

Be an Early Applicant Remote Hiring Remotely in USA 122K-162K Annually Mid level Remote Hiring Remotely in USA 122K-162K Annually Mid level Provide senior-level escalation and troubleshooting for GPU/HPC infrastructure, diagnosing hardware, driver, and kernel issues. Perform root-cause analysis across clusters, build automations and docs, mentor peers, collaborate with engineering for permanent fixes, and participate in on-call rotation and high-volume deployments. The summary above was generated by AI

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda’s mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.

If you’d like to build the world’s best AI cloud, join us., This is a salaried exempt role. The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.

About Lambda

  • Founded in 2012, with 500+ employees, and growing fast
  • Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove
  • We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG
  • Our values are publicly available: https://lambda.ai/careers
  • We offer generous cash & equity compensation
  • Health, dental, and vision coverage for you and your dependents
  • Wellness and commuter stipends for select roles
  • 401k Plan with 2% company match (USA employees)
  • Flexible paid time off plan that we all actually use, 122K-162K Annually Mid level 122K-162K Annually Mid level Software Provide senior-level technical escalation and support for GPU/HPC cloud infrastructure. Troubleshoot hardware, drivers, kernel, networking, and workload issues; perform root-cause analysis across clusters; build automations and documentation; mentor junior engineers; participate in on-call rotations and incident response; collaborate with engineering to implement permanent fixes. Top Skills: AnsibleCi/CdCudaDatadogDockerFirewallsGpudirect RdmaGrafanaInfinibandKubernetesLinuxNcclNvidia GpusNvlinkPrometheusRoceSlurmTcp/IpTerraformVpn NVIDIA, 108K-207K Annually Senior level 108K-207K Annually Senior level Artificial Intelligence * Computer Vision * Hardware * Robotics * Metaverse Provide onsite and remote technical support for NVIDIA Ethernet and AI infrastructure, troubleshoot Linux-based systems, debug networking and interoperability issues, collaborate with engineering and marketing, document support methodologies, and act as primary customer contact, spending at least one week per month onsite. Top Skills: AnsibleBashBgpChatgptCopilotCursorDockerEvpnGeminiGleanKubernetesLinuxNvidia Ethernet SwitchingNvidia Spectrum-XOspfPythonQosRoceTcpdumpVxlanWiresharkYaml

What you need to know about the Colorado Tech Scene

With a business-friendly climate and research universities like CU Boulder and Colorado State, Colorado has made a name for itself as a startup ecosystem. The state boasts a skilled workforce and high quality of life thanks to its affordable housing, vibrant cultural scene and unparalleled opportunities for outdoor recreation. Colorado is also home to the National Renewable Energy Laboratory, helping cement its status as a hub for renewable energy innovation.

Key Facts About Colorado Tech

  • Number of Tech Workers: 260,000; 8.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Lockheed Martin, Century Link, Comcast, BAE Systems, Level 3
  • Key Industries: Software, artificial intelligence, aerospace, e-commerce, fintech, healthtech
  • Funding Landscape: $4.9 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Access Venture Partners, Ridgeline Ventures, Techstars, Blackhorn Ventures
  • Research Centers and Universities: Colorado School of Mines, University of Colorado Boulder, University of Denver, Colorado State University, Mesa Laboratory, Space Science Institute, National Center for Atmospheric Research, National Renewable Energy Laboratory, Gottlieb Institute

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on jobs.ashbyhq.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

1:12 min

Addressing the competitive landscape of specialized hardware demands

Hazal Mestci +1 · Coffee With Developers

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

Videos

See all

Related articles

See all