> Markdown version of [/jobs/ext/2705777-principal-engineer-ai-platform-operations-in-california-city](https://www.wearedevelopers.com/jobs/ext/2705777-principal-engineer-ai-platform-operations-in-california-city). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Engineer - AI Platform & Operations in California City - **Company:** Energy Jobline - **Location:** California City, CA, United States (Remote available) - **Experience:** Experienced - **Salary:** $168,000.0 - $230,000.0 - **Contract:** Permanent contract - **Skills:** A/B Testing, Artificial Intelligence, Amazon Web Services, Microsoft Azure, Continuous Integration, Python (Programming Language), Software Product Management, Cloud Services, Tensorflow, Management of Software Versions, Cloud Platform System, Large Language Models, AI Platforms, Kubernetes, Low Latency, Machine Learning Operations, Docker - **Published:** September 4, 2026 - **Apply:** https://www.energyjobline.com/job/principal-engineer-ai-platform-operations-california-city-31476461 ## About the Role * Senior Expertise: 12+ years of engineering experience, with at least 4 years at a Principal or Distinguished level. * ML Infrastructure Mastery: Expert-level knowledge of model serving (e.g., Triton, vLLM, Ray Serve) and deep experience with Kubernetes and cloud ecosystems (AWS/GCP/Azure). * The MLOps Toolkit: Proven experience with tools like MLflow, Weights & Biases, or Kubeflow. * Coding & Systems: High proficiency in Python and a "systems-thinking" approach to Docker, Helm, and containerization. * AI Specialization: Hands-on experience operating LLM inference at scale and a deep understanding of the trade-offs between throughput and latency. * Platform Mindset: A track record of building internal platforms that treat other engineers as the primary customer, drastically improving engineering velocity. ## Description We are looking for a Principal Engineer to serve as the technical expert for our AI Platform & Operations. This isn't just a role about maintaining infrastructure; it's about owning the technical vision for the foundational systems that will power every AI product team across the company. Remote Role - can be based in either Washington or California State. Requirements By building robust, self-serve internal developer platforms, you'll enable our engineers to deploy, monitor, and scale AI models with unprecedented efficiency and safety. Your work ensures that our "Agentic" future is reliable, cost-effective, and cutting-edge. What You'll Be Doing * Architect the Future: Define the long-term technical roadmap for our AI platform, covering everything from model serving and feature stores to experiment tracking and CI/CD for ML. * Set the Gold Standard: Establish engineering benchmarks for deployment, versioning, A/B testing, and automated rollbacks. * Optimize & Scale: Lead strategies for GPU/compute efficiency and cost optimization while managing the complexities of LLM inference (quantization, batching, and latency) at scale. * Enhance Observability: Design sophisticated monitoring and alerting systems specifically tailored for AI workloads in production. * Champion Reliability: Drive platform stability and partner with application teams to ensure our infrastructure meets evolving product needs. * Technical Leadership: Mentor Senior engineers, lead architecture reviews, and evaluate the next of cloud services and ML frameworks. ## Related Videos - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Machine learning in the browser with TensorFlowjs](https://www.wearedevelopers.com/videos/155-machine-learning-in-the-browser-with-tensorflowjs) - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [Microservices: how to get started with Spring Boot and Kubernetes](https://www.wearedevelopers.com/videos/242-microservices-how-to-get-started-with-spring-boot-and-kubernetes) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud)