Artificial Intelligence Engineer

JR Software Solutions
United States
5 days ago
Apply on www.thejobnetwork.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Compensation
$140,000.0 - $180,000.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Python (Programming Language) Strategies of Testing Large Language Models Generative AI

Requirements

  • Bachelor’s degree in computer science, data science, engineering, statistics, or a related field. \n

  • 8+ years in software or model testing and evaluation, including hands-on evaluation of generative AI, LLM, or RAG systems. \n

  • Working knowledge of the NIST AI RMF and its Generative AI Profile (NIST AI 600-1). \n

  • Strong Python for evaluation harnesses, scoring scripts, and analysis notebooks. \n

  • Experience with LLM evaluation metrics, including LLM-as-judge scoring and how to validate it., Because of federal client requirements, U.S. citizenship is required, and you must have lived in the United States for at least 3 of the last 5 years. You must be able to obtain and maintain a federal Public Trust.

Benefits & conditions

JR Software Solutions Inc. (JRSS) is hiring an AI Test & Evaluation Lead in the Washington, DC area (Alexandria, VA, hybrid) to lead independent testing of generative AI, LLM, retrieval-augmented generation (RAG), and machine learning systems for a federal civilian client.

\n

\n

This is an evaluation role, not a development role. We don’t build the systems we test. Our job is to answer one question with evidence: does this system perform to a measurable standard, and is it ready to deploy? You will set the test strategy, measures, datasets, and pass criteria, and your results must be repeatable by anyone who re-runs them.

\n

\n

If you have evaluated LLMs for groundedness and hallucination, know why an LLM-as-judge score needs checking, and can write a test plan that holds up to review, this is your role.

\n

\n

About JRSS

\n

JRSS is a certified Women-Owned Small Business (WOSB), Small Disadvantaged Business (SDB), and HUBZone company headquartered in Tampa, FL. We deliver Cloud, AI/ML, Big Data and BI Analytics, ERP, and IT program management to Federal Civilian agencies and Fortune 500 clients. Our AI Assurance practice gives agencies independent evidence that their AI systems perform, behave safely, and are ready to deploy.

\n

\n

What you’ll do

\n \n

  • Write the AI Test & Evaluation Plan, test strategies, evaluation methods, test plans, and test cases. \n

  • Choose measures that fit each system: accuracy, groundedness, citation fidelity, instruction following, refusal behavior, bias, consistency, latency, and stability. \n

  • Evaluate RAG retrieval and generation separately; score ML models on precision, recall, F1, AUC, and calibration. \n

  • Build mission-specific test sets from the client’s domain terminology and realistic scenarios. \n

  • Make every result repeatable: versioned datasets and prompts; recorded model versions, settings, and seeds; immutable run logs. \n

  • Check AI-assisted scoring against human review, and cross-check vendor scores (e.g., Azure AI Foundry) with independent methods such as RAGAS. \n

  • Work with red team, IV&V, and cloud engineers on a FedRAMP-authorized Azure Government testing platform. \n

  • Map test evidence to the NIST AI Risk Management Framework to support agency AI governance and ATO decisions. \n

  • Write results, findings, and deployment readiness reports, and brief technical staff and leadership. \n

\n

\n, * A track record of test strategies, plans, and reports that stand up to review. \n

  • You live in the DC / Northern Virginia / Maryland area and can be on site as directed. \n

\n

\n

Nice to have

\n \n

  • Azure AI Foundry evaluations, RAGAS, promptfoo, or DeepEval. \n

  • AI red teaming with PyRIT, Garak, or NIST Dioptra; MITRE ATLAS and the OWASP Top 10 for LLM Applications. \n

  • Federal IV&V, ATO, or NIST SP 800-53 experience. \n

  • Azure Government or other FedRAMP environments. \n

  • Master’s or PhD; Azure AI Engineer, ISTQB, CSTE, or CISSP certification. \n

  • An active Public Trust. \n

\n

\n, $140,000 - $180,000 per year, based on experience, certifications, and clearance · medical, dental, and vision coverage · 401(k) · paid time off and holidays · hybrid schedule · professional development and certification support.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.thejobnetwork.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:22 min

Evaluating advanced artificial intelligence platforms for daily recruitment

Rudi Bauer Rudi Bauer +1 · Cappuccino with HR

2:37 min

Tracing the evolution from early AI to generative AI

Mike Mike · World Congress 2025

1:53 min

Actionable steps to improve team testing strategies immediately

Luise Freese Luise Freese · World Congress 2025

3:32 min

Fundamentals and limitations of large language models

Krzystof Czieslak · LIVE

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

2:14 min

Introduction to generative AI and content warnings

Cheuk Ho · World Congress 2023

Videos

See all

Related articles

See all