> Markdown version of [/jobs/ext/1907230-applied-ai-manager-evaluation-measurement](https://www.wearedevelopers.com/jobs/ext/1907230-applied-ai-manager-evaluation-measurement). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Applied AI Manager-Evaluation & Measurement - **Company:** Leonardo DRS - **Location:** Fort Walton Beach, FL, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Signals Intelligence, Large Language Models, Model Validation, Information Technology, Operational Systems - **Published:** August 3, 2026 - **Apply:** https://www.juju.com/job/00000000gldcyy ## About the Role + 5+ years in data, ML, or test engineering (regardless of total years exp), including 3+ years in evaluation or model validation, with recent hands-on large-language-model evaluation work. + Demonstrated bi-directional evaluation rigor: an example where your evaluation stopped something from shipping, and one where it cleared something others doubted. Expect us to probe the methodology of each. + Built a measurement practice where none existed. The first baselines, the first pipelines, the first dashboard leadership actually used, and it kept producing after you handed it off. + Data and measurement instrumentation skills: pipelines, dashboards, record-keeping, baselines from source-of-record data. + A demonstrated record of professional integrity under pressure: you have told someone, especially seniors "the number doesn't support that claim" and held. + Bachelor's degree and 5 years experience or equivalent combination of both. Engineering, Computer Science, Data Science, or a related technical field is preferred + Five years of relevant experience (or an equivalent combination of experience and training that provides the required knowledge, skills, and abilities) + Proficient technical expertise with demonstrated application + Excellent interpersonal, leadership, negotiation, communication and writing skills + U.S. citizenship, with the ability to obtain and maintain a U.S. Government security clearance. Preferred Qualifications + Experience wiring metrics to operational systems of record (not survey- or estimate-based measurement). + Statistical literacy for probabilistic acceptance: error budgets, confidence bounds, claims that survive challenge. + Finance-adjacent exposure: cost accounting, benefits realization, or audit (a plus, not a requirement). + Defense cost-accounting context. + Active or prior security clearance. ## Description We've used AI before, now we treat it as a first-class citizen at Airborne & Intelligence Systems (AIS). We build and deploy AI across signals intelligence, electronic warfare, air combat training, and the mission systems that connect them. This work shows up on our floor, in the field, and on the company's P&L, while meeting the demanding standards of our defense and commercial customers. If you want your work and your talent to count - this is the place. Two questions follow every AI system we build: does it actually work, reliably, on real jobs and is it actually worth the money we say it is? You own both answers. You design the test every capability must pass before it can ship, and you build the measurement that turns results into honest numbers leadership can trust. You hold the line. Nothing ships until it clears your bar - no test, no launch - and you're deliberately independent from the person who builds and deploys, so the check is real. When there's pressure to over-claim a result, you're the one who holds the number honest. At the heart of this function is a proprietary platform, the encoded way this business does AI. It takes any employee from a rough idea to a governed, well-formed use case, checks it against the rules, and prepares a recommendation a human decision-maker rules on. You don't just use it; you also build it. The quality bar and measurement logic you design get embedded into the platform itself, so the standard travels with the system to every team that touches it. You encode the acceptance standard and the value math once, and every team inherits both. This is a high impact, founding individual-contributor seat within the technical core of AI Excellence function with meaningful end to end work ownership. You report directly to top leadership of the function giving you outsize visibility, opportunity for personal growth and direct line to impact. Your peers will be AI leads for Solutions & Delivery/ Infrastructure & Deployment Job Responsibilities + Define "good enough to ship." Design the evaluation for each capability - baseline, acceptance criteria, error budget, human-review points - documented before launch. Gate the launch. + Build the measurement. Stand up the pipelines and record keeping that track what each capability actually delivers. + Establish the truth. Connect each measurement directly to the source system that holds the ground truth. + Define what deployed systems must log. Set the trace and telemetry standard that keeps every shipped capability diagnosable and auditable. + Produce the numbers. Deliver the monthly and quarterly results - forecast versus actual, clean and without double-counting - that the function defends to Finance. + Keep the front door open. Run the record intake so every AI idea is captured, and nothing happens off the books. + Scope boundaries. You are not the claim owner, governance or reporting. You build the instrument that proves or disproves the claims in an auditable manner. + Guide, maintain and monitor the operational excellence program and activities, ensuring integration and coordination with the AI Excellence leadership + Manage opportunities identified for improvement using AI Excellence principles and methodologies. + Effectively communicate with team members, senior leaders and key stakeholders on the status, objectives, risks, and mitigation plans associated with projects, as well as ensure team members are aware of integrated project timelines. + Provide training in support of AI excellence objectives and the execution of chartered projects + Manage budget, cost, schedule and rate of return for AI excellence activities ## Related Videos - [Introduction to Responsible AI: Balancing Value and Risk](https://www.wearedevelopers.com/videos/1972-introduction-to-responsible-ai-balancing-value-and-risk) - [Cutting LLM Costs Without Cutting Quality: How to Beat Proprietary LLMs with Fine-Tuned Open Source](https://www.wearedevelopers.com/videos/100151-cutting-llm-costs-without-cutting-quality-how-to-beat-proprietary-llms-with-fine-tuned-open-source) - [Enterprise Linux as Container Images](https://www.wearedevelopers.com/videos/1610-enterprise-linux-as-container-images) - [The age of agency: When products start to think and act](https://www.wearedevelopers.com/videos/1925-the-age-of-agency-when-products-start-to-think-and-act) - [AI in Regulated Industry - Validating AI-Enabled Products with PLM and Digital Twins](https://www.wearedevelopers.com/videos/2065-ai-in-regulated-industry-validating-ai-enabled-products-with-plm-and-digital-twins) - [How AI Models Get Smarter](https://www.wearedevelopers.com/videos/1374-how-ai-models-get-smarter) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [How to start an AI project for a good cause and boost your career](https://www.wearedevelopers.com/magazine/15-how-to-start-an-ai-project-for-a-good-cause-and-boost-your-career)