> Markdown version of [/jobs/ext/2898371-remote](https://www.wearedevelopers.com/jobs/ext/2898371-remote). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Remote - **Company:** MAG 24 LLC - **Location:** New York, NY, United States (Remote available) - **Experience:** Experienced - **Salary:** $83,200.0 - $124,800.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Document Management Systems, Enterprise Content Management, Human-Computer Interaction, Reinforcement Learning, Model Validation, Software Version Control - **Published:** September 14, 2026 - **Apply:** https://www.careerjet.com/jobad/us7088adbd95fa70d2dbfd2482b4267725 ## About the Role * At least 3 years of professional experience in linguistics, instructional design, technical writing, content design, or a closely related field * Direct experience developing or refining guidelines and rubrics for human evaluators in generative AI, RLHF, or model-assessment programmes * Demonstrated ability to resolve ambiguity and contradiction in complex written specifications * Experience translating specialist requirements into clear and practical instructions * Ability to work effectively across multiple subject-matter domains * A portfolio or concrete examples showing measurable improvements to guidelines, rubrics, or instructional materials * Demonstrable professional growth and increasing responsibility * Reliable availability for at least 35 hours per week during weekdays Educational Background * A degree in linguistics, instructional design, education, communications, technical writing, language studies, or a related field is highly relevant * Graduate-level education in applied linguistics, learning design, human-computer interaction, or information design may be helpful * Equivalent professional experience in AI evaluation, technical documentation, or guideline development may also be considered * Training in assessment design, taxonomy development, content strategy, or quality assurance may be valuable Nice to Have * Experience supporting large language model evaluation, reinforcement learning from human feedback, or AI training-data programmes * Familiarity with annotation platforms, human-feedback workflows, and rater calibration processes * Experience developing domain-specific guidance for finance, insurance, retail, legal, sports, or comparable fields * Knowledge of controlled language, information architecture, taxonomy design, or content governance * Experience conducting guideline usability tests or analysing inter-rater consistency * Familiarity with version control, documentation systems, and structured authoring tools * Previous collaboration with researchers, programme managers, engineers, and subject matter experts ## Description We are sharing a specialised full-time consulting opportunity for US-based linguists, instructional designers, and technical writers experienced in developing clear evaluation guidelines, structured rubrics, and human-rating instructions for generative AI programmes. This role supports a high-impact generative AI initiative focused on translating complex and potentially ambiguous programme requirements into precise, practical guidance for human evaluators. Selected professionals will develop rater-ready instructions across domains such as finance, retail, insurance, legal, and sports while resolving contradictions, defining edge cases, and improving consistency throughout evaluation workflows., Rater Guideline Development * Translate programme requirements into clear, structured, and actionable instructions for human evaluators * Develop guidelines that can be applied consistently across standard scenarios and complex edge cases * Define terminology, rating criteria, decision rules, exceptions, and escalation pathways * Ensure instructions are accessible to raters while preserving necessary domain-specific precision Rubric & Evaluation Framework Design * Design detailed scoring rubrics for evaluating generative AI outputs * Establish measurable criteria covering correctness, relevance, reasoning quality, completeness, and instruction adherence * Create examples and counterexamples illustrating different performance levels * Align evaluation frameworks with programme objectives and quality standards Ambiguity & Consistency Review * Review draft specifications for ambiguity, contradiction, missing information, and inconsistent terminology * Identify instructions that may lead to conflicting interpretations across raters * Revise guideline sets until they can be applied reliably with minimal escalation * Document concrete before-and-after improvements to written requirements and evaluation instructions Cross-Domain Instructional Translation * Convert specifications from finance, retail, insurance, legal, sports, and other specialist domains into rater-ready guidance * Collaborate with subject matter experts to understand domain-specific terminology and professional judgment * Preserve important technical nuance while making instructions clear to non-specialist evaluators * Maintain consistent structure and quality across multiple domain-specific guideline sets ## Related Videos - [Introduction to Responsible AI: Balancing Value and Risk](https://www.wearedevelopers.com/videos/1972-introduction-to-responsible-ai-balancing-value-and-risk) - [Carl Lapierre - Exploring Advanced Patterns in Retrieval-Augmented Generation](https://www.wearedevelopers.com/videos/1235-carl-lapierre-exploring-advanced-patterns-in-retrieval-augmented-generation) - [Edit Your Future: Queerverse Radical AI](https://www.wearedevelopers.com/videos/909-edit-your-future-queerverse-radical-ai) - [On the straight and narrow path - How to get cars to drive themselves using reinforcement learning and trajectory optimization](https://www.wearedevelopers.com/videos/205-on-the-straight-and-narrow-path-how-to-get-cars-to-drive-themselves-using-reinforcement-learning-and-trajectory-optimization) - [Adding knowledge to open-source LLMs](https://www.wearedevelopers.com/videos/1522-adding-knowledge-to-open-source-llms) - [Inside the Mind of an LLM](https://www.wearedevelopers.com/videos/1617-inside-the-mind-of-an-llm) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [13 AI Tools You Have to Try](https://www.wearedevelopers.com/magazine/219-13-ai-tools-you-have-to-try)