> Markdown version of [/jobs/ext/1894342-lead-ai-ml-engineer-remote-nationwide-or-hybrid-in-mn-dc](https://www.wearedevelopers.com/jobs/ext/1894342-lead-ai-ml-engineer-remote-nationwide-or-hybrid-in-mn-dc). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead AI/ML Engineer- Remote Nationwide or Hybrid in MN/DC - **Company:** Optum, Inc - **Location:** Eden Prairie, MN, United States (Remote available) - **Experience:** Expert - **Salary:** $145,500.0 - $249,500.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Audit Trail, Automation of Tests, Microsoft Azure, Cloud Engineering, Continuous Integration, Fraud Prevention and Detection, Python (Programming Language), Key Management, Machine Learning, OAuth, Regression Testing, Search Technologies, Twilio, Core Voice Platform, Azure Service Bus, Pytorch, Large Language Models, Multi-Agent Systems, Spring-boot, Rate Limiting, AI Platforms, Scikit Learn, Information Technology, HuggingFace, Cosmos DB, Machine Learning Operations, Api Design, Dynatrace, Serverless Computing, Microservices - **Published:** July 31, 2026 - **Apply:** https://jobs.localjobnetwork.com/apply/add/87881703/1 ## About the Role * Bachelor's degree or higher in Computer Science, Engineering, or related field, or equivalent professional experience * 8+ years building production software, including 3+ years as a technical lead for AI or ML product development (ownership from design through launch) * 5+ years professional Python experience and hands on use of ML or NLP libraries such as PyTorch, Hugging Face, or scikit learn, including deploying models or pipelines to production * 3+ years hands on building and operating Spring Boot microservices in production, including API design, automated testing, CI/CD, and on call or incident participation * Agentic and LLM systems experience: Proven shipped at least one LLM based system to production that uses tool calling or function calling to interact with external services or enterprise APIs, with defined evaluation and monitoring * RAG and vector search experience: Proven built and shipped retrieval augmented generation or semantic search solutions using vector search, embeddings, and external knowledge integration (for example Azure Cognitive Search or comparable tooling) * Azure architecture experience: Proven built and deployed production workloads on Azure, using multiple services such as Azure OpenAI, Azure Functions, Event Hubs, Cognitive Search, and Cosmos DB, with security, observability, and cost considerations * Required Qualification: Must be authorized to work in the United States without the need for current or future employer-sponsored visa sponsorship (e.g., H-1B, TN, F-1/OPT, CPT, or other employment-based visa status). Preferred Qualifications: * LLM evaluation and regression testing: Experience building automated evaluation harnesses for LLM and agent workflows, including golden datasets, offline and online testing, and measurable quality metrics (for example task success rate, groundedness, or human review agreement) * Responsible AI and adversarial testing: Hands on experience with prompt injection and data exfiltration testing, safety reviews, and implementing guardrails to reduce hallucinations and unsafe outputs in production * Production observability for agents: Proven implemented end to end monitoring for agentic systems, including distributed tracing, tool call success rates, latency and error budgets, and token and cost telemetry with actionable alerting * Security for tool calling and AI systems: Proven designed secure patterns for tool enabled agents, including least privilege access, secrets management, and policy based controls for tool/API execution (for example OAuth scopes, managed identity, and audit logging) * Platform scale and efficiency: Proven ability to optimize LLM or voice system performance and cost using techniques such as caching, batching, streaming responses, rate limiting, model routing, and fallback strategies ## Description * Lead AI System Design: Architect and evolve our multi-agent AI platforms, enabling agents to reason, plan, and interact with external tools via LLMs and modular service layers * Technical Ownership: Define standards, best practices, and technical vision for AI orchestration across product lines including voice agents, fraud detection, and agentic workflows * Multi-Agent Frameworks: Guide adoption of frameworks like LangChain, AutoGen, and Semantic Kernel; integrate emerging protocols such as Model Context Protocol (MCP) to scale tool and agent interoperability * AI Interface Innovation: Lead the design of agentic user experiences, enabling LLMs to act as intelligent interfaces to enterprise tools and APIs * Voice AI Strategy: Architect full-stack voice agent pipelines - from ASR and multi-turn dialogue to TTS and telephony integrations (SIP, Twilio, etc.) * ML & Fraud Systems: Oversee development and deployment of ML models for fraud and anomaly detection, emphasizing scalability, explainability, and real-time responsiveness * Cloud-Native Engineering: Lead AI/ML system deployment using Azure OpenAI, Azure Functions, Event Hubs, Cognitive Search, Cosmos DB, and other cloud-native tools * Mentorship & Delivery: Guide senior and junior engineers, lead architecture reviews, and drive cross-team technical delivery in a globally distributed environment * AI Governance & MLOps: Set standards for experimentation, monitoring, CI/CD pipelines, and lifecycle management of LLM and ML models You'llbe rewarded and recognized for your performance in an environment that will challenge you and give you clear direction on what it takes to succeed in your role as well asprovidedevelopment for other roles you may be interested in. ## Related Videos - [How to Avoid LLM Pitfalls - Mete Atamel and Guillaume Laforge](https://www.wearedevelopers.com/videos/1328-how-to-avoid-llm-pitfalls-mete-atamel-and-guillaume-laforge) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [20 billion requests a week: Upgrading Twilio's API gateway at scale](https://www.wearedevelopers.com/videos/100234-20-billion-requests-a-week-upgrading-twilio-s-api-gateway-at-scale) - [Keeping applications secure by evolving OAuth 2.0 and OpenID Connect](https://www.wearedevelopers.com/videos/100152-keeping-applications-secure-by-evolving-oauth-2-0-and-openid-connect) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) - [Minimal infrastructure for Real‑Time Phone Agents: transcripts in, responses out](https://www.wearedevelopers.com/videos/1736-minimal-infrastructure-for-real-time-phone-agents-transcripts-in-responses-out) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it)