Remote Data Annotation Jobs Madrid
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
Remote Data Annotation Jobs Madrid (Full-Time, Remote) Join production-grade AI/ML training workflows by labeling and evaluating data used to improve model performance across NLP, computer vision, and content safety systems.You will follow strict rubrics and maintain strong annotation guidelines compliance to protect training data quality for LLM training pipelines.Key Responsibilities Execute structured annotation tasks: classification, ranking, span labeling, and named entity recognition (NER)Perform LLM evaluation and prompt evaluation (helpfulness, factuality, groundedness, instruction-following)Contribute to RLHF workflows via preference judgments, rationale writing, and comparative response evaluationRun QA evaluation through self-checks, peer reviews, and audit responsesDocument edge cases, track error patterns, and propose guideline clarifications to reduce varianceProjects and Task Types Multilingual NER and text labelingContent safety labeling aligned to policy taxonomiesComputer vision annotation (bounding boxes, segmentation, image-text alignment) as neededCalibration sets and inter-review agreement improvement activitiesRequired Qualifications Experience with guideline-based labeling, evaluation frameworks, QA, data operations, or similar rubric-driven workAbility to interpret complex guidelines and apply consistent decisions at scaleComfort with web-based annotation tools, task queues, and documentation of edge casesStrong written English for rationale-based RLHF comparisons and clear communicationBenefits Competitive hourly rate: $30-$50/hrFully remote, full-time structured workflows with calibration and quality checkpointsWork on real-world LLM evaluation and multimodal annotation programs#J-*****-Ljbffr
Requirements
Remote Data Annotation Jobs Madrid (Full-Time, Remote) Join production-grade AI/ML training workflows by labeling and evaluating data used to improve model performance across NLP, computer vision, and content safety systems. You will follow strict rubrics and maintain strong annotation guidelines compliance to protect training data quality for LLM training pipelines.Key Responsibilities Execute structured annotation tasks: classification, ranking, span labeling, and named entity recognition (NER)Perform LLM evaluation and prompt evaluation (helpfulness, factuality, groundedness, instruction-following)Contribute to RLHF workflows via preference judgments, rationale writing, and comparative response evaluationRun QA evaluation through self-checks, peer reviews, and audit responsesDocument edge cases, track error patterns, and propose guideline clarifications to reduce varianceProjects and Task Types Multilingual NER and text labelingContent safety labeling aligned to policy taxonomiesComputer vision annotation (bounding boxes, segmentation, image-text alignment) as neededCalibration sets and inter-review agreement improvement activitiesRequired Qualifications Experience with guideline-based labeling, evaluation frameworks, QA, data operations, or similar rubric-driven workAbility to interpret complex guidelines and apply consistent decisions at scaleComfort with web-based annotation tools, task queues, and documentation of edge casesStrong written English for rationale-based RLHF comparisons and clear communicationBenefits Competitive hourly rate: $30-$50/hrFully remote, full-time structured workflows with calibration and quality checkpointsWork on real-world LLM evaluation and multimodal annotation programs#J-*****-Ljbffr
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.buscojobs.com.esGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Spanish Business Culture and Etiquette
What Are Large Language Models?
Dev Digest 121 - AI goes offline
Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production