> Markdown version of [/jobs/ext/737795-principal-engineer-ai-platform-infrastructure](https://www.wearedevelopers.com/jobs/ext/737795-principal-engineer-ai-platform-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Engineer, AI Platform & Infrastructure - **Company:** SpreeAI Corporation - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, Amazon Web Services, Microsoft Azure, Cloud Computing, Computer Clusters, Nvidia CUDA, Continuous Integration, Data Governance, Software Debugging, Distributed Systems, Monitoring of Systems, Python (Programming Language), Azure Machine Learning, Software Engineering, Management of Software Versions, Graphics Processing Unit (GPU), Google Cloud, Pytorch, Delivery Pipeline, Large Language Models, Model Validation, Generative AI, AI Platforms, Kubernetes, Low Latency, Optimization Algorithms, Deployment Automation, ONNX (Open Neural Network Exchange) Format, Machine Learning Operations, TensorRT, Hardware Infrastructure, Data Pipelines, Docker - **Published:** June 29, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=afc210036ddeb195 ## About the Role * 10+ years of software engineering / infrastructure experience, with 5+ years in ML infrastructure, MLOps, distributed systems, or AI platform engineering. * Deep experience with Python, PyTorch, Kubernetes, Docker, cloud infrastructure, and GPU-based workloads. * Strong understanding of distributed systems and large-scale ML infrastructure design. * Experience with ML workflow orchestration systems such as Ray, Kubeflow, Argo, Airflow, Flyte, or Metaflow. * Experience deploying and managing production inference systems using platforms like Triton, vLLM, TensorRT-LLM, Ray Serve, KServe, Seldon, BentoML, TorchServe, or custom services. * Strong understanding of inference optimization techniques such as batching, quantization, CUDA graphs, and memory-aware scheduling. * Experience with model registries, experiment tracking, CI/CD for ML, canary deployments, shadow traffic, rollback strategies, and production monitoring. * Strong cloud experience across AWS, GCP, Azure, or GPU-focused providers like CoreWeave, Lambda Labs, or RunPod. * Ability to debug performance bottlenecks across distributed systems, containers, networking, GPU memory, and storage layers. Strong ownership mindset with the ability to define architecture, set platform standards, and drive execution across teams. Nice to Have * Experience with multimodal, vision, or generative AI systems. * Experience with large-scale GPU clusters e.g. A100/H100, NCCL, and high-throughput data pipelines. * Experience designing evaluation and monitoring systems for generative AI workloads. * Familiarity with ML security, privacy, and data governance practices. * Experience building internal developer platforms for research teams. ## Description We are looking for a Principal Engineer to build the infrastructure, deployment pipelines, and observability systems that enable multimodal AI models to move from research prototypes to reliable, production-grade deployments powering real-time virtual try-on experiences for global retail partners. This role spans ML platform engineering, deployment systems, GPU infrastructure, and observability. You will partner closely with Applied Science, AI Platform, Product, and Partner Engineering to enable rapid research iteration and reliable model delivery at scale. What You'll Own ML Platform & Training Enablement * Build and operate SPREEAI's end-to-end ML platform spanning training, evaluation, deployment, and monitoring. * Enable scalable and reliable training workflows through orchestration, infrastructure, and resource management systems. * Define platform standards for model packaging, model registry, dataset lineage, experiment tracking, checkpointing, and deployment automation. Deployment, Inference & Observability * Enable reliable and scalable inference deployments through standardized serving, orchestration, and monitoring frameworks. * Build and operate model deployment pipelines with versioning, reproducibility, rollback, approval gates, evaluation gates, and production observability. * Establish production SLOs for latency, availability, error rate, GPU saturation, cold-start time, cost per inference, and model quality drift. * Standardize and support serving infrastructure using modern inference runtimes such as vLLM, NVIDIA Triton, TensorRT-LLM, Ray Serve, TorchServe, ONNX Runtime, or equivalent systems. GPU Infrastructure & System Efficiency * Design and manage GPU allocation, scheduling, and resource utilization across training and inference workloads. * Improve GPU utilization, throughput, latency, reliability, and cost efficiency across model lifecycle systems. * Design and operate model evaluation and benchmarking systems, including automated regression detection and quality gates for production releases. * Partner with research teams to productionize new capabilities by providing robust infrastructure, tooling, and deployment pathways., * Within 6 months, you will: * Create reliable research-to-production pathways for SPREEAI's core AI models. * Reduce manual model deployment friction through standardized pipelines and tooling. * Improve GPU utilization and reduce training and inference costs. * Establish robust observability and evaluation gates for production model releases. * Accelerate the delivery of new AI capabilities into partner-facing experiences. Why This Role Matters ## Related Videos - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)