Senior Platform Engineer

STN, inc.
United States
about 2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Cloud Engineering Computer Clusters Configuration Management Computer Programming Image Management Web Portals Python (Programming Language) Open Source Technology Cloud Services Software Engineering AI Infrastructure
+9 more
Istio Multi-Agent Systems Core Api Build Management Kubernetes Information Technology Linkerd (Service Mesh) Slurm Hardware Infrastructure

Job description

The Senior Platform Engineer builds and operates the multi-tenant orchestration, scheduling, and customer-facing platform layer that turns raw GPU infrastructure into a usable cloud service. This role is the software backbone of GPU One (GPUaaS)., * Design and build the orchestration layer (Kubernetes, Slurm, Run:ai, or comparable)

  • Manage multi-tenant isolation including namespaces, networking, storage, and quotas
  • Build customer-facing platform APIs, CLIs, web portals, and SDKs
  • Implement and operate image management, GPU operator, and node provisioning automation
  • Drive infrastructure-as-code and automation across the platform stack
  • Partner with SRE on platform reliability, SLO definition, and observability
  • Support TAM and Support engineers on customer-impacting platform issues
  • Maintain customer environment templates, configuration management, and rollout tooling
  • Participate in architecture review, design discussions, and technical roadmap
  • Drive continuous platform improvement and reduce operational toil

Requirements

Do you have experience in Software engineering?, Do you have a Bachelor’s degree?, * 6+ years in platform engineering, SRE, or cloud engineering at scale

  • Deep Kubernetes expertise including CRDs, operators, and multi-tenant patterns
  • Strong programming skills in Go, Python, or both
  • Experience operating GPU clusters or AI infrastructure at production scale
  • Bachelor’s degree in computer science or equivalent experience, * Experience with NVIDIA GPU Operator, MIG, MPS, and NCCL operator patterns
  • Familiarity with Slurm operator, Run:ai, KubeRay, or comparable AI orchestration
  • Service mesh experience (Istio, Linkerd) and multi-cluster networking
  • Open source contributions in the cloud-native or AI infrastructure ecosystem

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:42 min

Core API types for Angular signals

Daniela Bonvini · LIVE

2:53 min

Configuring dynamic proxy updates with Istio Pilot

Jan Mensch Jan Mensch · WWC Europe 2026

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · WWC Europe 2026

4:25 min

Exploring native core APIs and web standard support

Jo Franchetti Jo Franchetti · LIVE

7:15 min

Installing Istio programmatically with bash scripts

Thomas Südbröcker · LIVE

Videos

See all

Related articles

See all