> Markdown version of [/jobs/ext/2532267-principal-ai-architect-m365-ic3-team-intelligent-conversation-and-communications-cloud](https://www.wearedevelopers.com/jobs/ext/2532267-principal-ai-architect-m365-ic3-team-intelligent-conversation-and-communications-cloud). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal AI Architect - M365 IC3 Team (Intelligent Conversation and Communications Cloud) - **Company:** Microsoft - **Location:** Redmond, WA, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Systems Engineering, Automated Storage and Retrieval Systems, Cloud Computing, Computer Engineering, Continuous Integration, Data Visualization, Recommender Systems, Data Streaming, Enterprise Search, Delivery Pipeline, Large Language Models, Generative AI, Indexer, Information Technology, Build Tools, Machine Learning Operations, Data Pipelines - **Published:** August 20, 2026 - **Apply:** https://www.careerjet.com/job/usfd5ecbab620f2785c1c1d550298a344f/eaa ## About the Role experimentation systems, and engineering workflows so evaluation becomes a standard part of how products are built and shipped. Make evaluation results easy to access, interpret, and act on through dashboards, scorecards, quality reports, and product-health views. Build systems that connect product telemetry, offline evaluation, human judgment, automated evals, experimentation, RAG quality, agent behavior, and customer-quality signals. Understand and evaluate enterprise search, RAG, grounding, indexing, ranking, permissions, freshness, and relevance systems for products such as Copilot. Translate ambiguous product goals into measurable evaluation strategies, success criteria, timelines, and technical plans. Drive architecture decisions across components, services, data pipelines, model interfaces, search systems, retrieval layers, evaluation harnesses, dashboards, and reporting systems. Work with product leaders to prioritize evaluation investments and align them with product milestones, Statistics, Econometrics, Computer Science, Electrical or Computer Engineering, or related field, AND 4+ years related experience (e.g., statistics, predictive analytics, research) OR Doctorate in Statistics, Econometrics, Computer Science, Electrical or Computer Engineering, or related field AND 3+ years related experience (e.g., statistics, predictive analytics, research) Master's Degree in Statistics, Econometrics, Computer Science, Electrical or Computer Engineering, or related field AND 9+ years related experience (e.g., statistics, predictive analytics, research) OR Doctorate in Statistics, Econometrics, Computer Science, Electrical or Computer Engineering, or related field AND 6+ years related experience (e.g., statistics, predictive analytics, research) OR equivalent experience. 5+ years experience creating publications (e.g., patents, libraries, peer-reviewed academic papers). 2+ years experience presenting at conferences or other events in the outside research/industry community as an invited speaker. 5+ years experience conducting research as part of a research program (in academic or industry settings). 3+ years experience developing and deploying live production systems, as part of a product team. 3+ years experience developing and deploying products or systems at multiple points in the product cycle from ideation to shipping. Demonstrated experience evaluating LLMs, including designing eval datasets, defining quality metrics, analyzing model behavior, identifying failure modes, and using results to guide product or system improvements. Understanding of modern AI systems, including LLMs, agentic systems, RAG, ML pipelines, evaluation methodology, experimentation, and product telemetry. Ability to architect complex systems where multiple components, services, models, data flows, tools, retrieval systems, and product surfaces interact. Experience evaluating or building nondeterministic AI systems where quality must be understood statistically, behaviorally, and through product impact. Experience bringing ML, DS, LLM, or applied science concepts into production systems and product codebases. Ability to integrate evaluation into engineering systems such as CI/CD, build pipelines, release gates, dashboards, and monitoring workflows. Understanding of search, retrieval, grounding, relevance, ranking, and enterprise RAG concepts. Ability to define technical strategy, product-quality metrics, milestones, and execution plans across teams. Coding and technical design skills, with the ability to work directly in product codebases when needed. Effective communication skills with the ability to influence engineers, scientists, product managers, and executives. Track record of leading ambiguous, cross-functional technical initiatives from concept through delivery. Experience evaluating LLM-powered products, agents, enterprise search, recommendation systems, or generative AI applications. Experience with offline evals, online experimentation, human evaluation, red teaming, synthetic data, model monitoring, RAG evaluation, and agent behavior analysis. Experience with evaluation dashboards, scorecards, quality reporting, product-health monitoring, or data visualization systems. Familiarity with responsible AI, safety, reliability, privacy, security, permissions, compliance, and enterprise-readiness considerations for AI systems. Experience building evaluation platforms, experimentation systems, model observability, agent evaluation infrastructure, or product-quality infrastructure. Experience operating at principal, architect, or senior technical leadership level. ## Description and release decisions. Mentor senior engineers and applied scientists on building reliable, scalable, and reusable evaluation infrastructure. Stay current with LLM evaluation methods, agentic systems, RAG evaluation, benchmark design, prompt/model behavior, experimentation, and responsible AI practices. In this role, you will help evaluate products before the code is fully ready, before launch, and after they ship. You will work across product, engineering, applied science, and data science teams to bring rigorous LLM, RAG, agent, and AI evaluation practices into the product lifecycle. You will help IC3 and partner teams understand whether AI systems are working as intended, where they fail, how they improve, and what it takes to ship them responsibly at scale., Bachelor's Degree in Statistics, Econometrics, Computer Science, Electrical or Computer Engineering, or related field AND 6+ years related experience (e.g., statistics, predictive analytics, research) OR Master's Degree in ## Related Videos - [Optimizing Discovery: PostgreSQL's Role in Transforming GetYourGuide's Search](https://www.wearedevelopers.com/videos/1647-optimizing-discovery-postgresql-s-role-in-transforming-getyourguide-s-search) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) - [Your imaginations is (no longer) the limit: how Generative AI empowers people to be creative](https://www.wearedevelopers.com/videos/741-your-imaginations-is-no-longer-the-limit-how-generative-ai-empowers-people-to-be-creative) - [Inside the AI Revolution: How Microsoft is Empowering the World to Achieve More](https://www.wearedevelopers.com/videos/869-inside-the-ai-revolution-how-microsoft-is-empowering-the-world-to-achieve-more) - [Should we build Generative AI into our existing software?](https://www.wearedevelopers.com/videos/1129-should-we-build-generative-ai-into-our-existing-software) - [Dynamic Entities in .NET: Building Low-Code Systems on Top of Entity Framework Core](https://www.wearedevelopers.com/videos/100218-dynamic-entities-in-net-building-low-code-systems-on-top-of-entity-framework-core) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)