> Markdown version of [/jobs/ext/2582033-ai-ml-engineering-managers-lead-the-data-and-ml-engineering](https://www.wearedevelopers.com/jobs/ext/2582033-ai-ml-engineering-managers-lead-the-data-and-ml-engineering). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI/ML Engineering Managers lead the Data and ML Engineering - **Company:** Bain & Company - **Location:** Atlanta, GA, United States - **Experience:** Expert - **Salary:** $148,500.0 - $162,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Automated Storage and Retrieval Systems, Code Review, Information Engineering, Identity and Access Management, Python (Programming Language), Machine Learning, Operational Databases, Azure Machine Learning, Management of Software Versions, Large Language Models, Model Validation, Technical Debt, Generative AI, Pytest, Information Technology, Low Latency, Production Code, Machine Learning Operations, Terraform, Docker, Databricks - **Published:** August 9, 2026 - **Apply:** https://www.dice.com/job-detail/b559a8fe-f6b8-4261-97f3-9fe46a33c3d1 ## About the Role * Bachelor's degree in Computer Science, Engineering, Machine Learning, Data Science, Statistics, or a related field (or equivalent practical experience). * 8+ years of experience building and operating production data and ML systems, including model deployment, serving, and post-deployment monitoring. * 3+ years of direct people-management experience leading data, ML, or platform engineers, including hiring, performance management, and career development. * Demonstrated experience owning delivery for a team: roadmap planning, capacity allocation, and accountability for commitments and operational health. * Demonstrated experience owning the model and prompt lifecycle end-to-end (packaging, registry, promotion, rollout, monitoring, and rollback). * Demonstrated experience building and operating production RAG or retrieval systems end-to-end, from embedding and retrieval through re-ranking and evaluation. * Experience setting architectural direction across data and ML systems and leading design review for a team of engineers. * Experience partnering with Data Science, Product, and platform leadership to translate strategy into staffed, sequenced engineering work. * Strong Python for production ML, with the technical depth to contribute code and review critical pull requests while managing. * Experience working in a modern cloud ML platform environment (Databricks or AWS), including managed training / serving and governed model access. ## Description Senior AI/ML Engineering Managers lead the Data and ML Engineering function for the Diligence Platform: they own the people, delivery, and technical direction of the teams that build feature and embedding pipelines, model serving and LLMOps infrastructure, and production RAG and retrieval systems. You manage and develop a team of Data and ML Engineers - hiring, coaching, performance management, and career growth - while owning the roadmap, delivery commitments, and operational health of everything the team ships. You remain hands-on where it matters most: setting architectural direction, prototyping hard problems, reviewing critical code, and stepping into individual contribution when delivery, incidents, or complexity demand it. You partner with Data Science, the Agent / AI squad, Product, and Platform leadership to translate platform strategy into staffed, sequenced, and measurable engineering work. This is a player-coach role: the team's output, reliability, and engineering standards are the measure of success., * Manage a team of Data and ML Engineers: hiring and onboarding, goal setting, coaching, performance management, and career development. * Own delivery for the team's roadmap: scope and sequence work, set realistic commitments, track progress, and remove blockers before they become escalations. * Set and enforce engineering standards across data and ML systems - testing, typing, observability, runbooks, and SLAs - and hold the team to production quality from the first commit. * Own the operational health of the team's systems: on-call and incident response ownership, post-incident reviews, and follow-through on reliability and cost improvements. * Allocate capacity across pipeline, serving, retrieval, and evaluation workstreams; balance feature delivery against platform and technical-debt investment. * Partner with Data Science, the Agent / AI squad, Product, and Platform leadership to translate platform strategy into staffed, sequenced engineering work with clear ownership. * Grow senior engineers into technical leaders: delegate architectural ownership, coach through design review, and build bench depth behind every critical system. * Communicate roadmap, trade-offs, risks, and delivery status to platform leadership and non-technical stakeholders. Hands-On Engineering and Technical Direction (35%) * Set architectural direction for the platform's data, serving, and retrieval systems; own design review and the build-versus-buy calls that follow. * Contribute directly to production code where depth or delivery pressure requires it: inference and serving systems, model and prompt lifecycle in MLflow, LLMOps tooling, and RAG and retrieval pipelines. * Prototype hard or ambiguous problems personally to de-risk them before handing an established pattern to the team. * Own the evaluation and monitoring bar for production ML: golden datasets, metric definitions, regression gates in CI, and drift and hallucination monitoring. * Review critical pull requests and lead incident response on the highest-severity production ML and data issues. Other (15%): * Contribute to platform-wide engineering standards, repository conventions, and interviewing and hiring loops beyond the immediate team. * Mentor senior and mid-level engineers and Data Scientists on production ML, LLMOps, and data engineering practice. * Set the team's standards for AI-assisted development: how coding assistants are used, and how generated code is reviewed before it enters the codebase. * Use LLMs to generate first-draft documentation, runbooks, and status and evaluation reports; validate and refine outputs before publishing., * Builds and develops high-performing engineering teams: hires well, sets clear expectations, gives direct feedback, and manages underperformance early. * Delivery management: decomposes ambiguous platform goals into sequenced, owned, and estimable engineering work. * Coaches through design and code review rather than taking over; delegates architectural ownership deliberately. * Operates as the accountable owner for team systems: on-call structure, incident process, post-incident follow-through, and cost and reliability targets. * Communicates clearly upward and across: trade-offs, risk, and delivery status framed for both engineers and non-technical stakeholders. ML engineering / LLMOps * Strong Python: ML and serving code written to production engineering standards (type hints, Pydantic, pytest, Ruff, mypy strict). * MLflow: experiment tracking, model registry, custom model flavours, promotion workflows, and model serving configuration. * LLMOps tooling: prompt and instruction versioning, model gateways (e.g., Portkey), inference orchestration frameworks (LangChain, LlamaIndex, or equivalent), and response caching. * Model serving and inference optimisation: request batching, concurrency and throughput tuning, latency budgeting, and awareness of quantisation and hardware trade-offs. * RAG pipeline engineering: chunking strategies, contextual embedding, hybrid retrieval, cross-encoder re-ranking, and context assembly. * Data engineering breadth: batch and streaming pipelines, orchestration, data quality, and feature and embedding pipeline design at scale. * Model evaluation and monitoring: golden datasets, metric definition, calibration, LLM-as-judge, regression gates in CI, and production drift monitoring. * Docker and Kubernetes, and infrastructure familiarity: able to review ML-serving infrastructure (IAM roles, model endpoints, GPU / CPU workloads) provisioned via Terraform. Generative AI and agentic systems * Owns the team's inference and retrieval services that feed agent workflows as structured tool responses consumed by the Agent Gateway. * Sets the RAG quality bar: embedding and retrieval strategy, and recall / precision and groundedness benchmarks the team is held to. * Directs use of LLM-as-judge patterns in evaluation pipelines where qualitative criteria are required. * Builds feedback-to-evaluation and feedback-to-training-data loops from production signals, and staffs the work to sustain them. General * Treats every ML system as a production system from the first commit: tests, observability, a runbook, and an SLA are not optional - and holds the team to the same bar. * Raises model quality, reliability, and delivery risks proactively; does not wait for users or leadership to surface them. * Sets the standard for AI-assisted development: uses AI tooling to move faster, but reviews all generated code and documentation critically before it enters the codebase. * Strong communication skills; able to explain retrieval-quality, latency, cost, and staffing trade-offs to engineers, senior stakeholders, and non-technical audiences. * This role follows a hybrid model, requiring in-office presence at least 1 day per week ## Related Videos - [Navigating the AI Revolution in Software Development](https://www.wearedevelopers.com/videos/1266-navigating-the-ai-revolution-in-software-development) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [pytest: Simple, rapid and fun testing with Python](https://www.wearedevelopers.com/videos/213-pytest-simple-rapid-and-fun-testing-with-python) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Implementing Feature Environments with AWS and Terraform](https://www.wearedevelopers.com/videos/531-implementing-feature-environments-with-aws-and-terraform) - [Automagic Configuration in Python](https://www.wearedevelopers.com/videos/363-automagic-configuration-in-python) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts) - [From developer to manager – what does it take to become an engineering manager?](https://www.wearedevelopers.com/magazine/42-from-developer-to-manager-what-does-it-take-to-become-an-engineering-manager) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)