> Markdown version of [/jobs/ext/1201255-ai-red-team-engineer](https://www.wearedevelopers.com/jobs/ext/1201255-ai-red-team-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Red Team Engineer - **Company:** White Circle - **Location:** France (Remote available) - **Salary:** €60,000.0 - €90,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Software System Penetration Testing, Automation of Tests, Burp Suite, Software as a Service, Information Leak Prevention, Python (Programming Language), Open Web Application Security, Regression Testing, Red Team (Cyber Security), Web Applications, Scripting, Cloud Platform System, Chatbots, Postman, Large Language Models, Software Security, Pytest, Bug Reporting, Playwright, Api Design - **Published:** July 8, 2026 - **Apply:** https://fr.indeed.com/viewjob?jk=fb97826c46acc1ff ## About the Role * Experience with Burp Suite, Postman, Playwright, pytest. * Experience with modern LLM red-teaming automated agents and pipelines. * Familiarity with LangChain, LangGraph, LlamaIndex, RAG pipelines, AI agents, tool/function calling, and LLM-as-judge evaluation. * Familiarity with OWASP LLM Top 10, OWASP Web Top 10, MITRE ATLAS, or other AI security taxonomies. * Experience testing RAG systems, AI agents, tool-calling workflows, browser agents, or internal copilots. * Experience writing customer-facing security reports. * Experience with trust & safety, abuse prevention, fraud, moderation, or platform security. * Experience building eval pipelines, regression suites, dashboards, or CI-friendly security tests. * A track record in CTFs, red-team competitions, or responsible-disclosure / bounty programs. ## Description * Red-team LLM-powered systems: chatbots, copilots, RAG pipelines, AI agents, tool-calling workflows, and API-based AI products. * Test for jailbreaks, prompt injection, system-prompt and tool leakage, sensitive-data and context leakage, unsafe outputs, policy bypass, tool misuse, excessive agency, resource and token-cost abuse, and business-logic abuse. * Write lightweight Python to automate attacks, run prompt sets, call model APIs, collect and score responses, and generate repeatable reports. * Build and maintain an internal attack library: prompts, scenarios, test cases, regression tests, scoring rubrics, and reusable demo cases. * Turn model failures into clear reports: what happened, why it matters, how to reproduce it, how severe it is, and how to fix it. * Convert successful attacks into regression tests and product requirements. * Track new red-team and safety techniques and fold the useful ones into our tests. * Support GTM by producing strong, credible evidence for customer demos, security reviews, and sales conversations. You'll fit right in if you: * Genuinely love breaking things and reasoning adversarially. * Have a background in QA automation, AppSec, API/security/pen testing, or bug bounty. * Have strong Python scripting skills. * Have experience testing APIs, web apps, backends, or SaaS products. * Are hands-on with LLMs, prompts, system instructions, RAG, agents, and tool/function calling. * Understand LLM-specific abuse vectors (prompt injection, jailbreaks, data leakage, tool misuse, excessive agency, token-cost exhaustion). * Can find bypasses, abuse edge cases, chain failures, and reason about real-world impact. * Can separate real customer risk from low-impact prompt tricks. * Write clear, reproducible bug reports in clear English. * Can move fast without perfect requirements. * Hold a firm ethical line: you red-team to make systems safer, operate within scope and the law, and don't produce or traffic in genuinely harmful material. ## Related Videos - [pytest: Simple, rapid and fun testing with Python](https://www.wearedevelopers.com/videos/213-pytest-simple-rapid-and-fun-testing-with-python) - [Chatbots are going to destroy infrastructures and your cloud bills](https://www.wearedevelopers.com/videos/1130-chatbots-are-going-to-destroy-infrastructures-and-your-cloud-bills) - [Agentic employees in world's most downloaded FinTech app](https://www.wearedevelopers.com/videos/100123-agentic-employees-in-world-s-most-downloaded-fintech-app) - [Testing AI Agents: Automated Evaluation for Chatbots & RAG Systems](https://www.wearedevelopers.com/videos/100300-testing-ai-agents-automated-evaluation-for-chatbots-rag-systems) - [Automagic Configuration in Python](https://www.wearedevelopers.com/videos/363-automagic-configuration-in-python) - [The AI-Native Engineering Org: What’s Real, What’s Hype, What’s Next](https://www.wearedevelopers.com/videos/100004-the-ai-native-engineering-org-what-s-real-what-s-hype-what-s-next) ## Related Articles - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [Dev Digest 134 - Where pixels sing?](https://www.wearedevelopers.com/magazine/477-dev-digest-134-where-pixels-sing) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 132 - Binging WADFlix?](https://www.wearedevelopers.com/magazine/473-dev-digest-132-binging-wadflix) - [Dev Digest 122 - Cracks in the polyfill](https://www.wearedevelopers.com/magazine/457-dev-digest-122-cracks-in-the-polyfill)