Platform Engineer

Saicon Consultants Inc.
San Jose, CA, United States
3 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Computer Clusters Computer Networks Computer Engineering DevOps Monitoring of Systems Remote Direct Memory Access Tensorflow Prometheus Software Engineering AI Infrastructure
+8 more
Pytorch Grafana Multi-Cloud Kubernetes Infrastructure Automation Frameworks Information Technology Slurm Terraform

Job description

  • Build and extend platform capabilities to enable different classes of workloads (e.g., Large-scale AI training, inferencing etc).
  • Design and operate scalable orchestration systems using Kubernetes across both on-prem and multi-cloud environments.
  • Develop platform features such as pre-flight health checks, job status monitoring and post-mortem analysis.
  • Partner with development teams to extend the GPU developer platform with features, APIs, templates, and self-service workflows that streamline job orchestration and environment management.
  • Apply expertise in storage and networking to design and integrate CSI drivers, persistent volumes, and network policies that enable high-performance GPU workloads.
  • Production support on large-scale GPU clusters.

Requirements

We are seeking an AI Infrastructure / Platform Engineer to join our team building and operating large-scale GPU compute infrastructure that powers AI and ML workloads. The ideal candidate should be passionate about software engineering and possess leadership skills to independently deliver on multiple projects. They should be able to communicate effectively and work optimally with their peers within our larger organization., * Experience in Platform, Infrastructure, DevOps Engineering.

  • Deep hands-on experience with Kubernetes and container orchestration at scale.
  • Proven ability to design and deliver platform features that serve internal customers or developer teams
  • Experience building developer-facing platforms or internal developer portals (e.g. Custom workflow tooling)., * Hands-on experience in storage or network engineering within Kubernetes environments (e.g., CSI drivers, dynamic provisioning, CNI plugins, or network policy).
  • Experience with Infrastructure as Code tools like Terraform.
  • Background in HPC, Slurm, or GPU-based compute systems for ML/AI workloads.
  • Practical experience with monitoring and observability tools (Prometheus, Grafana, Loki, etc).
  • Understanding of machine learning frameworks (PyTorch, vLLM, SGLang, etc.).
  • High performance network and IB/RDMA tuning.

Academic Credentials:

  • Bachelor’s or master’s degree in computer science, computer engineering, electrical engineering, or equivalent.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:24 min

Comprehensive AI infrastructure stacks at the Linux Foundation

Matt White Matt White · WWC 2025

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · WWC Europe 2026

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · WWC 2025

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all