> Markdown version of [/jobs/ext/3040466-principal-ai-platform-engineer](https://www.wearedevelopers.com/jobs/ext/3040466-principal-ai-platform-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal AI Platform Engineer - **Company:** PEMCO MUTUAL INSURANCE COMPANY - **Location:** Seattle, WA, United States - **Experience:** Experienced - **Salary:** $250,000.0 - $300,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Microsoft Azure, Encodings, Nvidia CUDA, Information Leak Prevention, Routing, Performance Tuning, Role-Based Access Control, Regression Testing, Azure Machine Learning, System Availability, Large Language Models, Multi-Agent Systems, AI Platforms, Kubernetes, Machine Learning Operations, TensorRT, Hardware Infrastructure, Api Management - **Published:** September 23, 2026 - **Apply:** https://www.jofdav.com/jobs/59825101-principal-ai-platform-engineer ## About the Role 6+ years in platform engineering, SRE, or technical operations at senior/lead scope * 2+ years running LLM/AI systems in production (or 4+ years ML platform operations) * Hands-on open-weight model serving (vLLM, Triton/TensorRT-LLM or equivalent), including quantization and GPU right-sizing * LLM observability and evaluation experience (LangFuse, LangSmith, Arize class), or the demonstrated ability to stand it up * Experience building RAG systems: vector stores, embedding pipelines, document processing * FinOps/cost management for cloud consumption, ideally GPU or AI workloads * Working knowledge of AI security risks: prompt injection, data leakage, model abuse * Preferred: on-prem GPU infrastructure design, fine-tuning/LoRA, agent frameworks (MCP, LangChain/LangGraph, Semantic Kernel), regulated-industry experience ## Description Own AI economics and observability. Build the consumption telemetry for the entire AI estate (cost per agent, per workflow, per business outcome; token usage; GPU utilization) and the LLM observability layer under it (LangFuse-class tracing, evaluation, drift monitoring). Leadership decisions about AI spend run on your data. * Run the AI gateway. A single gateway fronting every provider and model (LiteLLM-class or Azure APIM GenAI): routing, fallback chains, quotas, and per-agent cost capture. This is the enforcement point for the economics. * Stand up open-weight serving. High-throughput serving of open-weight models on Azure GPU capacity (vLLM, Triton/TensorRT-LLM), quantization, right-sizing. You build the sizing evidence that justifies or kills any future on-prem investment. * Build the retrieval layer. Vector stores, embedding pipelines, and document processing for in-house AI builds, on a governed platform rather than one-off deployments. * Operate production agents. Deployment, monitoring, incident response, and retirement for the agent fleet (MCP-based orchestration, per-agent least-privilege identity). When an agent supporting a business workflow fails, you own recovery. * Shape the architecture. Represent platform reliability, security, and economics in AI solution reviews across teams; constructively challenge designs with data and propose alternatives. * Hold the governance line. RBAC for models and agents, prompt/output guardrails, a seat on the AI Governance Working Group with authority to block deployments that do not meet the bar. * Compute and acceleration stack: Azure GPU VMs and AKS GPU pools first; on-premise GPU build-out (hardware selection, CUDA stack, Kubernetes GPU scheduling) when the sizing data says so * Models and serving stack: open-weight model families with high-throughput serving (vLLM, NVIDIA Triton/TensorRT-LLM), quantization, fine-tuning and LoRA adaptation; managed frontier APIs (Azure OpenAI / AI Foundry) as the other half of the portfolio * Gateway and routing stack: a single AI gateway fronting every provider and model; routing and fallback chains, quotas, per-agent cost capture * Retrieval and data stack: vector stores, embedding and chunking pipelines, document processing, knowledge-source governance * Orchestration and agents stack: agent frameworks and protocols (MCP, LangChain/LangGraph, Semantic Kernel), tool registration, per-agent least-privilege identity * Observability, evaluation, governance stack: LangFuse-class tracing and cost telemetry, evaluation tooling with regression testing before prompt or model changes ship, guardrails, RBAC * Across all layers: model lifecycle from evaluation to retirement, and cost-against-capability optimization (model selection and routing, caching, batching, token budgets). ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [The Private AI Platform: Why Agentic Apps Need a Private Application Platform](https://www.wearedevelopers.com/videos/100162-the-private-ai-platform-why-agentic-apps-need-a-private-application-platform) - [A Brief History of Data Storage](https://www.wearedevelopers.com/videos/974-a-brief-history-of-data-storage) - [Creating a routing app with Google Maps API from scratch](https://www.wearedevelopers.com/videos/831-creating-a-routing-app-with-google-maps-api-from-scratch) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [JSON and Beyond](https://www.wearedevelopers.com/videos/968-json-and-beyond) ## Related Articles - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering)