Senior Solutions Engineer

TensorWave Inc.
Las Vegas, NV, United States
12 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Border Gateway Protocol Nvidia CUDA Software Debugging Linux Python (Programming Language) Remote Direct Memory Access Ansible Diagnostic Tools Reliability of Systems Kubernetes Hardware Infrastructure

Job description

We’re looking for a Senior Solutions Engineer to serve as the elite escalation point between our Global Operations Center (GOC) and our Core Engineering teams. You are the technical backstop for our most sophisticated customers teams training large models who cannot afford a single hour of downtime.

You’ll own the problems that go beyond runbooks, sitting at the intersection of customer success and engineering: resolving the hardest technical blockers and translating those findings into a more resilient product. If you’re an engineer who loves the detective work of kernel-level debugging and high-performance networking, and who also thrives in the high stakes environment of customer-facing resolution, this is your role.

What You’ll Do

  • Resolve Complex Escalations: Act as the final authority on issues exceeding GOC scope, utilizing code-level debugging and architectural investigation.
  • Direct Customer Engagement: Partner with customer technical leads to diagnose production issues, ensuring transparency and rapid resolution through active collaboration.
  • Iterative Problem Solving: Develop diagnostic scripts and workarounds to maintain customer operations while long-term patches are in development.
  • Drive Root Cause Analysis: Own end-to-end P1 resolution, partnering with TAMs to deliver clear, actionable post-incident analysis.
  • Bridge to Engineering: Convert recurring customer pain points into evidence-based feature requests, influencing product roadmap to resolve systemic failures.
  • Build Scalable Knowledge: Document non-obvious platform behaviors and refine GOC runbooks, ensuring institutional knowledge grows with every incident., All offers of employment are contingent upon verification of identity and authorization to work in the United States, as required by law.

Background Checks

Where permitted by law, employment may be contingent upon the successful completion of a job-related background check.

Data Privacy Notice

By submitting an application, you acknowledge that TensorWave may collect, use, and retain your personal information for recruiting and employment-related purposes in accordance with applicable data privacy laws.

Requirements

  • 5-9 years in Infrastructure Engineering, Platform Engineering, or SRE, with a specific focus on high-performance computing or large-scale AI stacks. Proven track record of managing complex production environments where system reliability is mission-critical.
  • Kubernetes Expert: Deep experience in cluster administration and scheduler internals; comfortable reading/modifying controller code.
  • AI/GPU Infrastructure Specialist: Proficient in orchestrating GPU workloads and diagnosing training job failures using ROCm or CUDA.
  • Network Pathologist: Skilled in RDMA/RoCEv2, SRIOV, and BGP; capable of interpreting switch telemetry to identify silent packet drops.
  • Linux Power User: Expert in kernel networking, hugepages, and cgroups; able to debug at the OS layer when applications are silent.
  • Builder Mindset: Proficient in Python and Ansible; capable of writing custom diagnostic tools to automate remediation.
  • Executive Communicator: Strong technical rigor when presenting findings to VPs of Engineering, maintaining trust while delivering difficult updates.

Preferred Qualifications

  • Prior experience in a customer-facing engineering role (e.g., Solutions Engineering, Technical Support Engineering).
  • Experience in high-uptime environments where 24/7/365 availability is required.

Benefits & conditions

Pulled from the full job description

  • Parental leave
  • 401(k)
  • Health insurance
  • Paid time off
  • Vision insurance
  • Health savings account
  • Dental insurance, * Stock Options
  • 100% paid Medical, Dental, and Vision insurance for Employees
  • Company Health Savings Account Contributions
  • 100% paid Short Term and Long Term Disability Insurance for Employees
  • Life and Voluntary Supplemental Insurance Options
  • Other Insurance Options, such as Pet & Legal Insurance
  • Various Supplementary Health Benefits, such as discounted Virtual Healthcare Appointments and Serious Illness Support
  • Flexible Spending Account
  • 401(k)
  • Employee Assistance Program
  • Flexible PTO
  • Paid Holidays
  • Parental Leave
  • Other In-Office Perks

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

1:42 min

Automating Skupper deployments using Ansible

Alex Soto Alex Soto · WWC 2024

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · WWC 2022

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · WWC Europe 2026

2:39 min

Experiencing core Linux capabilities for DevOps administration

Michael Cade · LIVE

Videos

See all

Related articles

See all