Senior/Staff SRE for AI/ML Platform Infrastructure

Xoriant Corporation
San Jose, CA, United States
27 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Bash Shell Nvidia CUDA Continuous Integration Distributed Computing Environment Distributed Systems Domain Name System (DNS) Github InfiniBand Python (Programming Language)
+17 more
Networking Basics Remote Direct Memory Access Prometheus Data Logging Google Cloud Load Balancing Cloud Platform System Grafana Firewalls (Computer Science) Gitlab-ci Kubernetes Machine Learning Operations Hardware Infrastructure Terraform Dynatrace Docker Jenkins

Requirements

  • Production on-call experience in a real rotation, with incident command and blameless postmortem practice.
  • Production Kubernetes and container experience (Docker), with cloud-native infrastructure patterns.
  • Hands-on production ownership on at least one major cloud (AWS, Google Cloud Platform, or Azure).
  • Terraform or OpenTofu proficiency.
  • Observability depth with Prometheus, Grafana, or equivalent for metrics, logging, and alerting, including dashboard and alert design.
  • Strong automation skills in Python, Bash, or Go.
  • Networking fundamentals: VPCs, load balancers, DNS, firewalls, cross-cloud connectivity.
  • CI/CD experience with GitHub Actions, GitLab CI, Jenkins, or ArgoCD.
  • Proven ability to troubleshoot complex distributed systems, largely self-directed.

Preferred Qualifications

  • GPU infrastructure and AI/ML workloads: Ray, Kubeflow, MLflow, or similar.
  • NVIDIA GPU orchestration: A100/H100 configuration, driver and CUDA runtime management.
  • Distributed training networking: RDMA, InfiniBand, EFA, NCCL.
  • Distributed tracing and OpenTelemetry instrumentation across services.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

1:02 min

Applying an ETL methodology to infrastructure configuration management

Axel Barbier · World Congress 2023

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

Videos

See all

Related articles

See all