Software Engineer - Platform Infrastructure (Rust, C++)

SPACEXAI LLC
Bellevue, WA, United States
27 days ago
Apply on www.techcareers.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$180,000.0
Working hours
Regular working hours

Tech stack

C++ (Programming Language) Profiling Computer Networks Software Debugging File Systems Distributed Systems Memory Management Linux Kernel Prometheus System Programming TCP/IP Wireshark
+9 more
Graphics Processing Unit (GPU) Istio Grafana Perf (Linux) Containerization Kubernetes Information Technology Optimization Algorithms Docker

Job description

  • Design, build, and implement a large-scale distributed system that powers one of the world’s largest supercomputing clusters.
  • Dive into the low-level stack to profile, debug, and optimize performance across diverse systems, including GPUs, Linux kernel, networking, and filesystems, to achieve peak efficiency.
  • Collaborate on hardware, software, and algorithm co-design to push the boundaries of AI training.
  • Maintain and innovate on our codebase to ensure scalability and reliability.
  • Develop tools to enhance team productivity and streamline workflows.

Requirements

All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates., * Systems programming experience in C, C++, or Rust

  • Computer systems fundamentals with a grasp of how computers execute code from transistors to high-level applications.
  • Hands-on expertise with Kubernetes (K8s), including cluster architecture, pod lifecycle, networking (CNI), storage (CSI), service mesh, and production-grade operations

PREFERRED SKILLS AND EXPERIENCE:

  • Collaborate in a fast-paced, open environment to design and foundational systems.
  • Strong debugging skills across the full stack - from kernel and OS up through container orchestration layers
  • Deep knowledge of operating systems internals (process scheduling, memory management, file systems, and synchronization primitives)
  • Proficiency in performance analysis, profiling, and low-level optimization techniques
  • Solid understanding of computer networks and the TCP/IP stack
  • Experience working with Linux kernel concepts or systems-level debugging tools (e.g., perf, gdb, strace, Wireshark)
  • Proficiency deploying and managing workloads using Kubernetes manifests, Helm, Operators, and GitOps workflows
  • Solid understanding of containerization technologies (Docker, containerd, crio) and their interaction with the Linux kernel
  • Experience with observability and monitoring in distributed systems (Prometheus, Grafana, VictoriaMetrics, OpenTelemetry, or similar)

Benefits & conditions

$180,000 - $440,000 USD

Base salary is just one part of our total rewards package at xAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.

SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice

About the company

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge.

Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity.

We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.techcareers.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:18 min

Prioritizing communication and structural awareness over strict tool mastery

Liam Hurrel +1 · World Congress 2021

5:02 min

Mapping distributed compute paradigms to modern vehicles

Joachim Werner · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

2:53 min

Configuring dynamic proxy updates with Istio Pilot

Jan Mensch Jan Mensch · World Congress 2026 Europe

3:50 min

Queues in TCP stacks and continuous network connections

Clemens Vasters Clemens Vasters · World Congress 2022

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

Videos

See all

Related articles

See all