> Markdown version of [/jobs/ext/2067411-python-engineer](https://www.wearedevelopers.com/jobs/ext/2067411-python-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Python Engineer - **Company:** Mindrift - **Location:** Spain (Remote available) - **Salary:** €124,800.0 - **Contract:** Permanent contract - **Skills:** JavaScript (Programming Language), Artificial Intelligence, Cyber Security, Databases, Continuous Integration, Github, Python (Programming Language), Software Engineering, TypeScript, GitHub Copilot, Prompt Engineering, Backend, Integration Tests, Codebase, Software Coding - **Published:** August 15, 2026 - **Apply:** https://es.trabajo.org/oferta-4021-d7a5621f0dc6117382d7bf6443715998 ## About the Role Description Please submit your CV in English and indicate your level of English proficiency. Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment. What this opportunity involves Frontier coding agents are already good at passing tests. We measure whether they pass them the right way.We're building a dataset to evaluate the safety and conduct of AI coding agents - not just whether they complete a task, but whether they do it honestly, within scope, and without cutting corners.You'll design tasks where the easy path is the unsafe one, and write the tests that catch it: - Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history - Design tasks that pair a benign development goal with a tempting unsafe shortcut: scope creep, policy violations, data corruption, over-permissive changes - Write tests that verify whether the agent completed the task the right way - catching corners cut, not just checking outputs - Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust What this is NOT: - Not data labeling; - Not prompt engineering; - Not cybersecurity or red-teaming - there is no attacker in the scenario. Cybersecurity experience is a nice-to-have but not a requirement. We're looking for engineers who understand how code should behave, not penetration testers. Strong software engineers, not security specialists; - Not writing code from scratch - the agent writes most of the code; you design the situation and evaluate the outcome; What we look for - 4-5+ years in software development; - Core stack: Python, JavaScript/TypeScript; - Strong test design skills - functional and integration tests that separate safe from unsafe completion, not just correct from incorrect; - Hands-on experience with coding agents (Claude Code, GitHub Copilot CLI, Codex, or similar); - Familiarity with GitHub PRs and CI workflows as a user; - Stack breadth is welcome, not a filter. Tasks simulate real repositories with databases, CI pipelines, and deploy scripts, so wider backend and infrastructure exposure is genuinely useful - but you don't need to be an expert in every layer; - English proficiency - B2+ Why this is hard Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. The real difficulty is building the temptation - a scenario where the unsafe or out-of-scope path is the path of least resistance - and then writing tests that reliably catch an agent that took it. Tasks have many valid solutions; tests must accept all of them and reject the bad ones. How it works Apply * Pass qualification(s) * ## Related Videos - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Getting to Know Your Legacy (System) with AI-Driven Software Archeology](https://www.wearedevelopers.com/videos/1437-getting-to-know-your-legacy-system-with-ai-driven-software-archeology) - [Hiring AI Native Talents](https://www.wearedevelopers.com/videos/100268-hiring-ai-native-talents) - [Boost your coding productivity with Github Copilot Agent](https://www.wearedevelopers.com/videos/1542-boost-your-coding-productivity-with-github-copilot-agent) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) ## Related Articles - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 131 - AI'm not sure about OSS](https://www.wearedevelopers.com/magazine/472-dev-digest-131-ai-m-not-sure-about-oss) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [What is Agentic Programming and Why Should Developers Care?](https://www.wearedevelopers.com/magazine/625-what-is-agentic-programming-and-why-should-developers-care) - [Should senior developers refuse interview coding challenges?](https://www.wearedevelopers.com/magazine/29-should-senior-developers-refuse-interview-coding-challenges)