Data Scientist
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+8 more
Job description
Since 2012, our client has helped mission-critical government organizations and businesses face their most daunting technology challenges. Their team has been a trusted partner to many government agencies and is extremely familiar with a wide variety of systems, policies, and procedures. Our client is also a distinguished custom software development firm dedicated to delivering premium solutions tailored for businesses and governmental needs. They are home to top-tier technology professionals recognized as industry pioneers, comprehensive engineers, and reliable consultants. These experts are adept at clear communication, excel in resolving complex challenges where others may falter, and are skilled in actualizing an organization’s vision.
The Data Scientist’s primary responsibility is to design, build, and maintain benchmarks for model testing and agentic workflow testing using InspectAI. This individual will define evaluation datasets, test scenarios, metrics, scoring approaches, and repeatable evaluation workflows that measure model and agent performance in a consistent and auditable manner.
The Data Scientist will also clean, transform, validate, and analyze supporting data, generate metrics and visualizations, and produce findings that support model evaluation, reporting, and decision-making. This individual will partner with Researchers to ground benchmark design in sound methodology and collaborate with Developers to operationalize evaluation workflows into stable, repeatable tooling.
The Data Scientist should also be able to document dataset lineage, scoring rationale, benchmark assumptions, and evaluation outputs so that results are reproducible, traceable, and usable within constrained or airgapped environments.
Responsibilities
Design, implement, and maintain AI model benchmark suites that measure model performance against defined evaluation criteria.
Curate evaluation datasets, develop scoring methods, and produce clear and repeatable results reporting.
Develop benchmarks and test scenarios that evaluate agentic workflows and multi-step task performance.
Measure tool use, task completion, reliability, and other performance indicators across agentic systems.
Curate, validate, and version evaluation datasets and test scenarios.
Develop adversarial and edge-case datasets designed to stress-test model and agent behavior.
Produce scoring frameworks, statistical analyses, visualizations, and written findings.
Translate raw evaluation results into clear and actionable information for technical and nontechnical stakeholders.
Document dataset lineage, scoring rationale, benchmark assumptions, evaluation methodology, and evaluation outputs.
Collaborate with Researchers and Developers to operationalize evaluation workflows into stable and repeatable tools.
Requirements
Experience with Python, InspectAI, MongoDB, and Jupyter notebooks. Experience designing or supporting AI model evaluation frameworks, LLM benchmarking, and agentic workflow evaluation. Experience evaluating tool use, multi-step task performance, task completion, and agent reliability. Strong knowledge of statistical analysis, experimental design, data cleaning, data validation, and data visualization. Experience curating, validating, and versioning evaluation datasets and test scenarios. Experience documenting dataset lineage, traceability, benchmark assumptions, scoring rationale, and evaluation results. Experience working within cloud-based or containerized environments. Experience using Git-based version control. Familiarity with Jira, Confluence, or similar collaboration and documentation tools. Ability to operate effectively within airgapped or otherwise constrained environments. Technologies and Skills InspectAI Python MongoDB Jupyter notebooks Data visualization tools Statistical analysis Data cleaning and validation LLM benchmarking Agentic workflow evaluation Experimental design Dataset lineage and traceability Cloud and containerized environments Git-based version control Jira Confluence Airgapped and constrained environments
Education Requirements
BS in Computer Science or a similar technical field. Four years of relevant experience may be substituted in lieu of a bachelor’s degree.
Clearance Requirements
Applicants selected will be subject to a security investigation and may need to meet eligibility requirements for access to classified information. An active TS/SCI with Full Scope Poly clearance is required. Please note that the Full Scope Poly currently needs to be held by the NSA or have been held within the past two years.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.clearancejobs.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How to Become an AI Engineer
Stephan Gillich - Bringing AI Everywhere
Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?
Navigating the AI Shift