> Markdown version of [/jobs/ext/1503945-llm-red-team-specialist-for-ai-model-evaluation](https://www.wearedevelopers.com/jobs/ext/1503945-llm-red-team-specialist-for-ai-model-evaluation). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # LLM Red Team Specialist for AI Model Evaluation - **Company:** Careerjet All Rights Reserved - **Location:** Washington, DC, United States (Remote available) - **Salary:** $124,800.0 - $187,200.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Data Analysis, Python (Programming Language), Machine Learning, Language Modeling, Red Team (Cyber Security), Large Language Models, Model Validation, Git, Machine Learning Operations - **Published:** July 30, 2026 - **Apply:** https://www.careerjet.com/job/register/usa9657b6421450f43179c04c8b04072e5 ## About the Role * MSc or PhD in a STEM field, or equivalent practical experience in a research-heavy role involving data analysis and coding. * At least 1 year of experience in research, research engineering, security, or AI evaluation. * Proven ability to identify vulnerabilities, edge cases, or failure modes in large language models or ML systems, through red teaming, adversarial testing, security research, or rigorous model evaluation. * Working proficiency in Python and Git, with the ability to script probes and analyses independently. * Strong familiarity with LLM capabilities, limitations, and common evaluation techniques. * Preferred experience in AI model training, model evaluation, or benchmark and task authoring. * High attention to detail, creativity in finding what others miss, strong written communication, and ability to work independently on ambiguous, open-ended problems. * Capacity to engage reliably for approximately 35 hours per week. Work Terms * Employment type, pay frequency, and compliance are handled via W-2 employment through an employer-of-record that administers payroll, benefits, and onboarding. ## Description Design and execute adversarial, multi-step probes that reveal where frontier language models appear competent but quietly fail. You will create short, focused challenge tasks, run experiments to reproduce failures, and collaborate closely with researchers to convert findings into robust evaluation benchmarks. Key Responsibilities * Probe models, exploring behavior on coding, machine learning, and analytic tasks to find subtle errors and failure modes. * Design challenge tasks that surface weaknesses, are difficult for models, and remain fair to grade. * Run reproducible experiments, capture evidence, and write clear, actionable findings that others can reproduce. * Work with task authors to close loopholes, eliminate shortcuts, and tighten grading criteria. * Share insights with researchers and colleagues to iteratively improve benchmarks and evaluations. ## Related Videos - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [Cutting LLM Costs Without Cutting Quality: How to Beat Proprietary LLMs with Fine-Tuned Open Source](https://www.wearedevelopers.com/videos/100151-cutting-llm-costs-without-cutting-quality-how-to-beat-proprietary-llms-with-fine-tuned-open-source) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Data Privacy in LLMs: Challenges and Best Practices](https://www.wearedevelopers.com/videos/1218-data-privacy-in-llms-challenges-and-best-practices) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 196: AI Killed DevOps, LLM Political Bias & AI Security](https://www.wearedevelopers.com/magazine/659-dev-digest-196-ai-killed-devops-llm-political-bias-ai-security) - [Who Owns Your Content in the Age of LLMs?](https://www.wearedevelopers.com/magazine/610-who-owns-your-content-in-the-age-of-llms) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Dev Digest 210: AI Agents Are Go! Is MCP Dead? LLMs Crack Anonymity](https://www.wearedevelopers.com/magazine/709-dev-digest-210-ai-agents-are-go-is-mcp-dead-llms-crack-anonymity)