> Markdown version of [/jobs/ext/1933296-platform-software-engineer](https://www.wearedevelopers.com/jobs/ext/1933296-platform-software-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Platform Software Engineer - **Company:** Oracle - **Location:** Redwood City, CA, United States (Remote available) - **Experience:** Experienced - **Salary:** $92,500.0 - $209,500.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, Nvidia CUDA, Distributed Systems, Python (Programming Language), Oracle Warehouse Builder, SAS (Software), Scientific Computating, Software Engineering, Data Streaming, System Software, Large Language Models, Multi-Cloud, Kubernetes, Free and Open-Source Software, Data Management, Slurm, Machine Learning Operations, TensorRT, Oracle Cloud Infrastructure - **Published:** August 5, 2026 - **Apply:** https://arc.dev/remote-jobs/j/redirect/p9zp0rsquk ## About the Role * 3-6 years of software engineering experience, ideally in distributed systems, data platforms, or ML infrastructure. * Strong Python; working knowledge of Go or Rust is a plus. * Hands-on experience with Kubernetes, container runtimes, and at least one workflow engine (Argo, Airflow, Prefect, Temporal). * Familiarity with GPU workloads, CUDA toolchains, or inference serving frameworks (vLLM, TensorRT-LLM, Triton). * Experience with LLM agent frameworks (LangGraph, CrewAI, or similar) is a strong plus. * Comfort working in ambiguity, shipping iteratively, and reasoning about topology - not just features. Nice to have * Prior work on multi-cloud or hybrid deployments. * Exposure to QEC, quantum-classical hybrid workloads, or scientific computing. * Open-source contributions to the AI infra ecosystem. ## Description We're hiring a Platform Software Engineer to help build the execution layer behind OCI's AI and GPU growth motion. You'll work across Oracle data platforms, GPU scheduling, orchestration pipelines, and MLOps agentic tooling - turning customer demand into production-grade systems that scale across our cluster fleet. This is a builder role on a small, high-leverage team. What you'll do * Integrate with Oracle data platforms (OCI Streaming, 23ai, Object Storage, Data Flow) to move training and inference data reliably across customer and internal pipelines. * Build and extend GPU schedulers and capacity-aware placement logic for A100, H100, H200, and Blackwell fleets. * Develop orchestration pipelines for training, fine-tuning, and inference workloads using Kubernetes, Argo, and Slurm where appropriate. * Ship MLOps agentic tooling - observability, automated triage, cost and SLO agents - that reduces operator load on large GPU deployments. * Partner with PMs, SAs, and customer-facing teams to convert field requirements into reusable components, not one-off scripts ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere)