> Markdown version of [/jobs/ext/2043181-remote](https://www.wearedevelopers.com/jobs/ext/2043181-remote). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Remote - **Company:** MAG 24 LLC - **Location:** New York, NY, United States (Remote available) - **Salary:** $93,600.0 - **Contract:** Contract - **Skills:** Artificial Intelligence, Information Systems, Data Sharing, Workflow Management Systems, Google Drive, Large Language Models, Model Validation, Stripe, Information Technology - **Published:** August 13, 2026 - **Apply:** https://www.careerjet.com/jobad/us77107e698c302a9642912da2a06f26b5 ## About the Role * Advanced practical experience using MCP, plugins, and AI connectors * Frequent use of connected LLM tools, ideally several times per week * Heavy personal use of AI for planning, research, organisation, and decision-making * An active LLM account with approximately 6 months or more of regular usage history * Experience using AI for high-context, multi-step personal workflows * Strong written judgment, reasoning, and attention to detail * Ability to explain clearly why an AI output is effective, incomplete, unsafe, or unrealistic * Experience designing and applying structured evaluation rubrics * Availability to contribute at least 20 hours per week * Ability to complete assigned tasks within approximately 24 hours * Current residence in the United States Educational Background * A degree in computer science, information systems, research, operations, behavioural science, communications, or a related discipline may be helpful * Formal technical credentials are secondary to extensive hands-on experience with LLM tools and connected workflows * Professional or project-based experience in AI evaluation, quality assurance, user research, or structured data review may strengthen an application * Equivalent practical expertise gained through sustained personal and professional AI usage may also be considered Nice to Have * More than 100 hours of prior rubric design, evaluation, or quality-assessment experience * Familiarity with Google Drive, Notion, travel platforms, and other connected applications * Experience evaluating AI assistants across personal planning and life-organisation tasks * Background in model evaluation, human data, user research, quality assurance, or AI training * Strong understanding of context management, personalisation, and tool-use failure modes * Experience documenting complex workflows through screen recording * Familiarity with health research, travel planning, dining decisions, home services, or career-search workflows * Experience identifying subtle issues involving model overreach, missing context, and unrealistic execution ## Description We are sharing a specialised part-time consulting opportunity for advanced LLM users with strong hands-on experience using Model Context Protocol tools, plugins, and connectors for complex personal workflows. This role supports an AI research initiative focused on evaluating how effectively AI assistants complete personalised, multi-step tasks using connected tools such as Google Drive, Notion, travel platforms, and other plugins or connectors. Selected professionals will create realistic tasks, execute workflows while recording their screens, assess model performance, and develop detailed evaluation rubrics grounded in practical everyday use., Personal Workflow Design * Create realistic prompts involving complex, high-context personal tasks * Develop scenarios across travel, health research, dining, activity planning, home services, career search, and personal organisation * Incorporate genuine preferences, constraints, trade-offs, and success criteria * Design tasks requiring planning, judgment, context retention, and connected-tool usage MCP, Plugin & Connector Testing * Use MCP-enabled tools, plugins, and connectors to complete multi-step workflows * Test integrations involving Google Drive, Notion, travel services, and similar platforms * Evaluate whether AI systems select and use connected tools appropriately * Identify failures involving permissions, context retrieval, sequencing, or action execution Screen-Recorded Task Execution * Complete assigned workflows while recording the screen * Clearly demonstrate the actions, tools, and decisions involved in each task * Document where the AI succeeds, overreaches, misses context, or produces impractical results * Complete tasks within required turnaround windows Model Evaluation & Rubric Development * Judge whether outputs are personalised, realistic, useful, safe, and well-reasoned * Write clear explanations of model strengths, weaknesses, and failure patterns * Create detailed scoring rubrics for complex personal-assistant tasks * Apply evaluation criteria consistently across model outputs * Identify incomplete reasoning, unrealistic recommendations, and incorrect tool use, * Help improve how AI assistants support complex real-world personal workflows * Evaluate advanced systems using practical connected tools and personal context * Influence how models handle planning, preferences, constraints, and multi-step actions * Apply deep LLM experience across travel, health, productivity, careers, and everyday decision-making * Build experience in MCP, connector evaluation, rubric development, and personalised AI research * Access potential ongoing work following successful completion of the initial trial period Contract Details * Independent contractor role * Fully remote within the United States * Expected commitment of at least 20 hours per week * Initial ramp-up period of approximately 1-2 days * Ability to complete assigned tasks within approximately 24 hours is required * Desktop or laptop computer required; Chromebooks are not supported * Screen recording is required during task execution * Candidates must be willing to sign a data-sharing consent form electronically * Initial trial period used to assess quality, consistency, and project fit * Competitive rates between $45-$185 per hour depending on expertise, task complexity, and project scope * Weekly payments via Stripe or Wise * Task availability may begin after an initial project setup period * Projects may be extended, shortened, or adjusted depending on scope and performance * Work will not involve access to confidential or proprietary information from any employer, client, or institution About the Platform This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams. By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: . ## Related Videos - [Navigating Growth, Scaling Challenges, and Office Expansions with David Singleton, CTO at Stripe](https://www.wearedevelopers.com/videos/100362-navigating-growth-scaling-challenges-and-office-expansions-with-david-singleton-cto-at-stripe) - [Smart City, Smart Mobility](https://www.wearedevelopers.com/videos/954-smart-city-smart-mobility) - [Hate organising your photos? Try it with 5 Terabytes](https://www.wearedevelopers.com/videos/79-hate-organising-your-photos-try-it-with-5-terabytes) - [How to Avoid LLM Pitfalls - Mete Atamel and Guillaume Laforge](https://www.wearedevelopers.com/videos/1328-how-to-avoid-llm-pitfalls-mete-atamel-and-guillaume-laforge) - [Open Source AI, To Foundation Models and Beyond](https://www.wearedevelopers.com/videos/1375-open-source-ai-to-foundation-models-and-beyond) - [Build Your Own Subscription-based Course Platform](https://www.wearedevelopers.com/videos/321-build-your-own-subscription-based-course-platform) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [13 AI Tools You Have to Try](https://www.wearedevelopers.com/magazine/219-13-ai-tools-you-have-to-try) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Who Owns Your Content in the Age of LLMs?](https://www.wearedevelopers.com/magazine/610-who-owns-your-content-in-the-age-of-llms) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How to Tell if Something Was Written by ChatGPT](https://www.wearedevelopers.com/magazine/314-how-to-tell-if-something-was-written-by-chatgpt)