> Markdown version of [/jobs/ext/2105876-machine-learning-engineer-lead](https://www.wearedevelopers.com/jobs/ext/2105876-machine-learning-engineer-lead). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Engineer Lead - **Company:** Compunnel Inc. - **Location:** Raleigh, NC, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Architectural Patterns, Microsoft Azure, Continuous Integration, Python (Programming Language), Google Cloud, Cloud Platform System, Large Language Models, Reliability of Systems, Generative AI, Containerization, AI Platforms, Kubernetes, Low Latency, Machine Learning Operations - **Published:** August 18, 2026 - **Apply:** https://www.dice.com/job-detail/eecb816e-dc7d-4a76-946e-73c26d8bafd4 ## About the Role reasoning-driven agentic AI systems. Design orchestration patterns for tool use, API invocation, and structured function calling. Lead the implementation and governance of Model Context Protocol (MCP) servers to standardize tool integration and context management. Define guardrails, permissions, security controls, and audit mechanisms for enterprise-safe AI systems. Establish and maintain best practices for MLOps, CI/CD, observability, scalability, and system reliability. Design and implement scalable inference systems using containerization and Kubernetes. Drive the deployment and optimization of LLM, Generative AI, and RAG solutions in production environments. Design cloud-based AI/ML architectures across AWS, Azure, or Google Cloud Platform. Establish technical standards and architectural patterns for AI/ML and agentic systems across engineering teams. Embed Responsible AI principles into platform architecture and engineering practices. Provide technical leadership, mentorship, and guidance to senior engineers and engineering teams. Collaborate with cross-functional teams to influence technical direction and ensure alignment with enterprise AI platform strategy. Support people management, leadership, and team development activities as required. Required Qualifications 10+ years of experience building and deploying production-grade machine learning systems at scale. Strong experience with LLMs, Generative AI, and RAG deployments in production environments. Strong Python development background. Expertise designing and implementing AI/ML systems in cloud environments such as AWS, Azure, or Google Cloud Platform. Hands-on experience with Kubernetes, containerization, and scalable inference systems. Experience designing agentic AI systems and tool orchestration frameworks. Experience implementing and governing MCP servers or structured architectures for tool integration and context management. Experience with large-scale distributed ML systems and enterprise platform engineering. Experience establishing MLOps, CI/CD, observability, and system reliability practices. Demonstrated people management, technical leadership, or mentorship experience. Strong understanding of high-availability and low-latency AI/ML architectures. Ability to define technical standards and influence architecture and engineering decisions across teams. Education: Bachelors Degree ## Related Videos - [One AI API to Power Them All](https://www.wearedevelopers.com/videos/1601-one-ai-api-to-power-them-all) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Swapping Low Latency Data Storage Under High Load](https://www.wearedevelopers.com/videos/746-swapping-low-latency-data-storage-under-high-load) - [The Cloud is Calling: Answer with In-Demand Skills](https://www.wearedevelopers.com/videos/945-the-cloud-is-calling-answer-with-in-demand-skills) - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/videos/100333-instant-kai-sandboxes-with-vcluster-multi-tenant-multi-scheduler-gpu-sharing) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Everything a Developer Needs to Know About MCP with Neo4j](https://www.wearedevelopers.com/magazine/604-everything-a-developer-needs-to-know-about-mcp-with-neo4j) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud)