> Markdown version of [/jobs/ext/2635830-senior-staff-sre-for-ai-ml-platform-infrastructure](https://www.wearedevelopers.com/jobs/ext/2635830-senior-staff-sre-for-ai-ml-platform-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior/Staff SRE for AI/ML Platform Infrastructure - **Company:** Xoriant Corporation - **Location:** San Jose, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Microsoft Azure, Bash Shell, Nvidia CUDA, Continuous Integration, Distributed Computing Environment, Distributed Systems, Domain Name System (DNS), Github, InfiniBand, Python (Programming Language), Networking Basics, Remote Direct Memory Access, Prometheus, Data Logging, Google Cloud, Load Balancing, Cloud Platform System, Grafana, Firewalls (Computer Science), Gitlab-ci, Kubernetes, Machine Learning Operations, Hardware Infrastructure, Terraform, Dynatrace, Docker, Jenkins - **Published:** August 10, 2026 - **Apply:** https://www.dice.com/job-detail/d744c2e3-27ee-4308-ba3f-33fd4014c682 ## About the Role * Production on-call experience in a real rotation, with incident command and blameless postmortem practice. * Production Kubernetes and container experience (Docker), with cloud-native infrastructure patterns. * Hands-on production ownership on at least one major cloud (AWS, Google Cloud Platform, or Azure). * Terraform or OpenTofu proficiency. * Observability depth with Prometheus, Grafana, or equivalent for metrics, logging, and alerting, including dashboard and alert design. * Strong automation skills in Python, Bash, or Go. * Networking fundamentals: VPCs, load balancers, DNS, firewalls, cross-cloud connectivity. * CI/CD experience with GitHub Actions, GitLab CI, Jenkins, or ArgoCD. * Proven ability to troubleshoot complex distributed systems, largely self-directed. Preferred Qualifications * GPU infrastructure and AI/ML workloads: Ray, Kubeflow, MLflow, or similar. * NVIDIA GPU orchestration: A100/H100 configuration, driver and CUDA runtime management. * Distributed training networking: RDMA, InfiniBand, EFA, NCCL. * Distributed tracing and OpenTelemetry instrumentation across services. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [The Road to MLOps: How Verivox Transitioned to AWS](https://www.wearedevelopers.com/videos/1050-the-road-to-mlops-how-verivox-transitioned-to-aws) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Our GitOps approach for deploying an Identity Provider and an API Gateway in a SaaS company](https://www.wearedevelopers.com/videos/776-our-gitops-approach-for-deploying-an-identity-provider-and-an-api-gateway-in-a-saas-company) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers)