Remote

MAG 24 LLC
New York, NY, United States
23 days ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Contract
Employment type
Part-time (≤ 32 hours)
Compensation
$93,600.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Information Systems Data Sharing Workflow Management Systems Google Drive Large Language Models Model Validation Stripe Information Technology

Job description

We are sharing a specialised part-time consulting opportunity for advanced LLM users with strong hands-on experience using Model Context Protocol tools, plugins, and connectors for complex personal workflows. This role supports an AI research initiative focused on evaluating how effectively AI assistants complete personalised, multi-step tasks using connected tools such as Google Drive, Notion, travel platforms, and other plugins or connectors. Selected professionals will create realistic tasks, execute workflows while recording their screens, assess model performance, and develop detailed evaluation rubrics grounded in practical everyday use., Personal Workflow Design

  • Create realistic prompts involving complex, high-context personal tasks
  • Develop scenarios across travel, health research, dining, activity planning, home services, career search, and personal organisation
  • Incorporate genuine preferences, constraints, trade-offs, and success criteria
  • Design tasks requiring planning, judgment, context retention, and connected-tool usage

MCP, Plugin & Connector Testing

  • Use MCP-enabled tools, plugins, and connectors to complete multi-step workflows
  • Test integrations involving Google Drive, Notion, travel services, and similar platforms
  • Evaluate whether AI systems select and use connected tools appropriately
  • Identify failures involving permissions, context retrieval, sequencing, or action execution

Screen-Recorded Task Execution

  • Complete assigned workflows while recording the screen
  • Clearly demonstrate the actions, tools, and decisions involved in each task
  • Document where the AI succeeds, overreaches, misses context, or produces impractical results
  • Complete tasks within required turnaround windows

Model Evaluation & Rubric Development

  • Judge whether outputs are personalised, realistic, useful, safe, and well-reasoned
  • Write clear explanations of model strengths, weaknesses, and failure patterns
  • Create detailed scoring rubrics for complex personal-assistant tasks
  • Apply evaluation criteria consistently across model outputs
  • Identify incomplete reasoning, unrealistic recommendations, and incorrect tool use, * Help improve how AI assistants support complex real-world personal workflows
  • Evaluate advanced systems using practical connected tools and personal context
  • Influence how models handle planning, preferences, constraints, and multi-step actions
  • Apply deep LLM experience across travel, health, productivity, careers, and everyday decision-making
  • Build experience in MCP, connector evaluation, rubric development, and personalised AI research
  • Access potential ongoing work following successful completion of the initial trial period

Contract Details

  • Independent contractor role
  • Fully remote within the United States
  • Expected commitment of at least 20 hours per week
  • Initial ramp-up period of approximately 1-2 days
  • Ability to complete assigned tasks within approximately 24 hours is required
  • Desktop or laptop computer required; Chromebooks are not supported
  • Screen recording is required during task execution
  • Candidates must be willing to sign a data-sharing consent form electronically
  • Initial trial period used to assess quality, consistency, and project fit
  • Competitive rates between $45-$185 per hour depending on expertise, task complexity, and project scope
  • Weekly payments via Stripe or Wise
  • Task availability may begin after an initial project setup period
  • Projects may be extended, shortened, or adjusted depending on scope and performance
  • Work will not involve access to confidential or proprietary information from any employer, client, or institution

About the Platform This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams. By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: .

Requirements

  • Advanced practical experience using MCP, plugins, and AI connectors
  • Frequent use of connected LLM tools, ideally several times per week
  • Heavy personal use of AI for planning, research, organisation, and decision-making
  • An active LLM account with approximately 6 months or more of regular usage history
  • Experience using AI for high-context, multi-step personal workflows
  • Strong written judgment, reasoning, and attention to detail
  • Ability to explain clearly why an AI output is effective, incomplete, unsafe, or unrealistic
  • Experience designing and applying structured evaluation rubrics
  • Availability to contribute at least 20 hours per week
  • Ability to complete assigned tasks within approximately 24 hours
  • Current residence in the United States

Educational Background

  • A degree in computer science, information systems, research, operations, behavioural science, communications, or a related discipline may be helpful
  • Formal technical credentials are secondary to extensive hands-on experience with LLM tools and connected workflows
  • Professional or project-based experience in AI evaluation, quality assurance, user research, or structured data review may strengthen an application
  • Equivalent practical expertise gained through sustained personal and professional AI usage may also be considered

Nice to Have

  • More than 100 hours of prior rubric design, evaluation, or quality-assessment experience
  • Familiarity with Google Drive, Notion, travel platforms, and other connected applications
  • Experience evaluating AI assistants across personal planning and life-organisation tasks
  • Background in model evaluation, human data, user research, quality assurance, or AI training
  • Strong understanding of context management, personalisation, and tool-use failure modes
  • Experience documenting complex workflows through screen recording
  • Familiarity with health research, travel planning, dining decisions, home services, or career-search workflows
  • Experience identifying subtle issues involving model overreach, missing context, and unrealistic execution

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

5:30 min

Building components of a real-world LLM lifecycle

Maxim Salnikov Maxim Salnikov · LIVE

3:54 min

Evaluating commercial cloud storage costs for large archives

Arto Liukkonen · LIVE

3:18 min

Utilizing AI and public data sharing for urban planning

Christian Wiegand +3 · World Congress 2024

10:22 min

Managing recurring subscriptions and products with Stripe

Dávid Lévai · LIVE

1:59 min

Building culturally aware LLMs for global audiences

Werner Vogels Werner Vogels +1 · World Congress 2026 Europe

3:16 min

Voice AI architecture and interactive kiosk demo

Lee Boonstra · LIVE

Videos

See all

Related articles

See all