Software Engineering Evaluation Specialist
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+13 more
Job description
- Invent a realistic developer scenario - a real bug, a broken ETL, a missing feature - not a toy problem.
- Build a reproducible Docker environment with pinned dependencies.
- Write a pytest that verifies outcomes, not specific commands - deterministic, non-flaky, and does not leak the fix.
- Write an instruction.md that reads like a Jira ticket a developer would receive.
- Write a reference solve.sh proving the task is solvable.
- Calibrate difficulty so current state-of-the-art agents solve the task 20-60% of the time.
- Iterate based on feedback from expert QA reviewers.
-
Later: review other authors’ tasks as a QA reviewer. Not in scope
- Data labeling, prompt engineering.
- Production code to ship - you design problems and verification for AI agents.
- Leetcode puzzles - scenarios must look like real developer work.
-
Not every candidate task ships - quality over quantity., Apply * Pass qualification (90-minute sample-task screen + short behavioral interview) * Join a project * Complete tasks * Get paid. Time commitment
- Onboarding: ~10 hours per first task.
- Steady state: ~5 hours per task, 2-4 parallel tasks per author.
- Realistic weekly load: 8-20 hours. Higher volume available for top performers.
- You choose when and how to contribute; tasks must be submitted by the deadline and meet acceptance criteria.
Requirements
- 3+ years of production software development in one backend stack - Python, Go, Node.js, Java, or Rust. Depth in one stack beats breadth.
- Python + pytest fluency - required regardless of primary stack. The task harness is pytest-based even when the broken app is in another language. Fixtures, parametrize, monkeypatch, timeouts, conftest.py.
- Docker authoring - reproducible Dockerfiles, pinned dependencies, multi-stage builds when needed, non-root user.
- Linux & Bash - comfort debugging inside containers (strace, lsof, journalctl); shell beyond set -euo pipefail.
- AI coding agent experience - Claude Code, Cursor, Roo Code, or similar, on non-trivial work. You can cite a specific time the AI was confidently wrong and how you caught it.
- English - B2+ written., + Domain depth in Security, System Administration (nginx / systemd / cron), Scientific Computing (NumPy / PyTorch / SciPy), DevOps, or Git internals.
Benefits & conditions
- Paid contributions, rates up to $35/hour*.
- Task-based compensation equivalent to hourly rate, depending on performance and volume.
- Some projects include incentive payments. *Rates vary based on expertise, skills assessment, location, project needs, and other factors. Higher rates may be provided to highly specialized experts. Lower rates may apply during onboarding or non-core project phases. Payment details are shared per project.
About the company
Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment., Analog Devices Spain, Valencia Software Engineering Payroll Specialist Banco Santander Sa España FS Banking Custom Software Engineering Specialist Accenture Madrid PhD Physicist AI Evaluation Specialist Mercor España Software Engineering AI Trainer (Spain) 45-80 per hour Anyone AI Madrid, Madrid Volver a la última búsqueda Trabajos ) Otros trabajos ) Software Engineering Evaluation Specialist ( volver a la última búsqueda Recibir ofertas similares por correo electrónico No gracias, llévame a la oferta de empleo Al crear una alerta, aceptas nuestros Términos y condiciones y Política de privacidad, y el uso de cookies. Inscribirse en esta oferta
Profesiones
- Técnico
- Recepciónista
- Administrador
- Ventas
- Enfermero
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Dev Digest 131 - AI'm not sure about OSS
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
Dev Digest 121 - AI goes offline
Dev Digest 137 - AI'm not sure about this