Remote

MAG 24 LLC
New York, NY, United States
about 2 months ago

Role details

Contract type
Permanent contract
Employment type
Part-time (≤ 32 hours)
Compensation
$41,600.0 - $62,400.0
Working hours
Regular working hours
Languages
English, Marathi
Job source

Tech stack

Artificial Intelligence Software System Penetration Testing Cyber Security Data Auditing Machine Learning Red Team (Cyber Security) Software Safety Reverse Engineering Model Validation Stripe Data Analytics

Job description

We are sharing a specialised part-time consulting opportunity for Marathi-English bilingual professionals experienced in AI safety evaluation, red team testing, adversarial review, vulnerability classification, and structured feedback on sensitive text-based AI outputs. This role supports current and upcoming remote consulting opportunities focused on AI safety evaluation, bilingual red team testing, conversational model assessment, misuse-risk review, vulnerability annotation, and high-quality project execution. Selected professionals will test AI systems using structured adversarial scenarios, identify safety weaknesses, classify risks, and produce clear English-language evaluation artifacts across English and Marathi contexts., Professionals in this role may contribute to: Bilingual AI Safety & Red Team Testing

  • Review English and Marathi AI outputs for safety, reliability, bias, misinformation, and harmful-behavior risks
  • Stress-test conversational AI models and agents using structured adversarial scenarios
  • Evaluate model behavior across multi-turn conversations, sensitive topics, and edge-case prompts
  • Identify vulnerabilities that require stronger safety controls, clearer refusals, or improved response quality

Vulnerability Classification & Risk Review

  • Annotate failures, classify vulnerabilities, and flag recurring safety patterns
  • Apply taxonomies, benchmarks, and project-specific playbooks to keep testing consistent
  • Assess misuse cases, bias exploitation, prompt-injection scenarios, and socio-technical risk patterns at a high level
  • Generate high-quality human evaluation data through careful review and structured judgment

Reproducible Documentation & Evaluation Artifacts

  • Produce clear reports, datasets, test cases, and written summaries that support model improvement
  • Document findings reproducibly so results can be reviewed, compared, and acted upon
  • Explain risks clearly for both technical and non-technical audiences
  • Maintain accuracy, consistency, and strong attention to detail across submitted evaluations, * Apply Marathi-English bilingual expertise to structured AI safety and red team evaluation work
  • Contribute to stronger, safer, and more reliable AI systems through careful adversarial testing
  • Work on flexible assignments aligned with language skills, safety judgment, and structured analysis
  • Build experience in human data-driven AI safety evaluation and bilingual risk review
  • Remote structure with competitive hourly compensation

Contract Details

  • Independent contractor role
  • Fully remote with flexible scheduling
  • Eligible professionals may be based in approved project locations depending on project needs
  • Native-level English and Marathi fluency are required for project work
  • Work is text-based and may involve sensitive topics such as bias, misinformation, harassment, or harmful-behavior risks
  • Topic areas will be communicated before exposure to content, and participation in higher-sensitivity projects may depend on candidate comfort and project fit
  • Part-time commitment depending on project availability
  • Competitive rates between $20-$30 per hour depending on expertise and project scope
  • Weekly payments via Stripe or Wise
  • Projects may be extended, shortened, or adjusted depending on scope and performance
  • Work will not involve access to confidential or proprietary information from any employer, client, or institution

About the Platform This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams. By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: .

Requirements

  • Native-level fluency in both English and Marathi
  • Prior experience in AI red teaming, adversarial testing, cybersecurity, trust and safety, socio-technical risk review, or conversational AI evaluation
  • Ability to think adversarially while staying structured, careful, and methodical
  • Experience using frameworks, benchmarks, or rubrics rather than unstructured testing alone
  • Strong written communication skills and ability to explain safety findings clearly
  • Comfort reviewing text-based content involving sensitive topics under clear guidelines
  • Adaptability across project types, safety categories, and evaluation workflows

Educational Background

  • Formal degree requirements may vary based on project needs
  • Backgrounds in AI safety, cybersecurity, linguistics, policy, trust and safety, social science, psychology, writing, data evaluation, or technical analysis may be highly relevant
  • Practical experience in red team testing, model evaluation, content risk analysis, or structured review work may also be valuable

Nice to Have

  • Experience with adversarial ML concepts, jailbreak datasets, prompt injection, RLHF/DPO attack patterns, or model behavior testing
  • Cybersecurity experience such as penetration testing, exploit analysis, reverse engineering, or security assessment
  • Socio-technical risk experience involving harassment, misinformation, abuse analysis, bias testing, or conversational AI safety
  • Creative probing background, including psychology, acting, writing, role-play design, or unconventional adversarial thinking
  • Experience producing reproducible reports, labeled datasets, structured risk notes, or benchmark-style evaluation artifacts

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on careerjet.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:18 min

Adapting engineering interviews to evaluate critical thinking and curiosity

Alex Laubscher Alex Laubscher +3 · WWC 2025

10:22 min

Managing recurring subscriptions and products with Stripe

Dávid Lévai · LIVE

3:16 min

Advantages of reproducible configurations and instantaneous rollbacks

Álvaro Martín Lozano · LIVE

4:11 min

Introduction to cloud-native application developer security

Micah Silverman · WWC 2022

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

2:10 min

Integrating Stripe checkout using serverless API routes

Christian K · JS Congress

Videos

See all

Related articles

See all