> Markdown version of [/jobs/ext/1481430-causal-reasoning-model-evaluation-specialist](https://www.wearedevelopers.com/jobs/ext/1481430-causal-reasoning-model-evaluation-specialist). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Causal Reasoning Model Evaluation Specialist - **Company:** Human Union Data, Inc. - **Location:** United States (Remote available) - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Model Validation - **Published:** July 29, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=24fb4e65893c2767 ## About the Role * Graduate-level training or equivalent applied experience in causal reasoning model evaluation research review or a closely related field for Causal Reasoning Model Evaluation Specialist work. * Hands-on experience publishing, teaching, or advising on the topic at a professional level. * Comfort applying multi-page rubrics consistently across long batches. * Clear written reasoning that cites methods, papers, or worked examples. * Reliable async availability for at least 10 hours per week., * PhD, postdoc, or industry research experience in the topic area. * Prior work reviewing AI-assisted research tooling and its failure modes. * Multilingual fluency for non-English papers and corpora., * Scientific reasoning * Method validation * Citation review * Quantitative analysis * Causal Reasoning Model Evaluation research review * Frontier evaluation * Rubric calibration * Failure analysis * Causal * Reasoning Work model Remote - US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US. ## Description Causal Reasoning Model Evaluation Specialist is a remote review track for evaluating AI outputs across causal reasoning model evaluation research review reasoning, calculations, and research workflows. Reviewers grade derivations and assumptions, reproduce key results, and document the correct method so the modeling team can train on it. Why this role matters Causal Reasoning Model Evaluation research review models live or die on whether their derivations actually hold up under scrutiny. AuraOne uses scientific specialists to grade outputs the way a peer reviewer would - checking assumptions, reproducing key steps, and capturing the right method alongside the wrong one., * Review AI outputs against current causal reasoning model evaluation research review methods, conventions, and prior work for Causal Reasoning Model Evaluation Specialist assignments. * Reproduce or sanity-check key derivations, calculations, or experimental claims. * Flag dimensional, methodological, and citation errors with structured severity tags. * Capture the corrected reasoning or worked example so the modeling team can train on it. * Adjudicate disputed answers against textbooks, papers, or community standards. * Maintain reviewer-quality scores in inter-rater calibration cycles., * Reproduce a causal reasoning model evaluation research review derivation from a model output and flag any algebraic or dimensional errors. * Grade a model's literature summary against the cited papers and rate the citation quality. * Adjudicate a disputed answer between two reviewers using textbook methods. * Audit a 25-row batch for rubric consistency and report drift to the program lead. ## Related Videos - [Introduction to Responsible AI: Balancing Value and Risk](https://www.wearedevelopers.com/videos/1972-introduction-to-responsible-ai-balancing-value-and-risk) - [The shadows of reasoning – new design paradigms for a gen AI world](https://www.wearedevelopers.com/videos/1000-the-shadows-of-reasoning-new-design-paradigms-for-a-gen-ai-world) - [Edit Your Future: Queerverse Radical AI](https://www.wearedevelopers.com/videos/909-edit-your-future-queerverse-radical-ai) - [AI is dead, long live AK](https://www.wearedevelopers.com/videos/1093-ai-is-dead-long-live-ak) - [Staying Safe in the AI Future](https://www.wearedevelopers.com/videos/521-staying-safe-in-the-ai-future) - [Building Trustworthy AI in Industry: Beyond Traditional Cybersecurity](https://www.wearedevelopers.com/videos/1948-building-trustworthy-ai-in-industry-beyond-traditional-cybersecurity) ## Related Articles - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Prompt Engineering is a Job of the Past](https://www.wearedevelopers.com/magazine/342-prompt-engineering-is-a-job-of-the-past) - [Why AI is Always Better When a Human is Using It](https://www.wearedevelopers.com/magazine/532-why-ai-is-always-better-when-a-human-is-using-it) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts)