Research Scientist - Computer Vision (Body Pose Detection)

MECKA ASSOC
New York, NY, United States
8 days ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$150,000.0 - $200,000.0
Working hours
Regular working hours
Job source

Tech stack

Agile Methodology Artificial Intelligence Computer Vision Computer Clusters Information Engineering Distributed Computing Environment Kinematics Motion Capture Rapid Prototyping Process Pytorch Deep Learning Legacy Systems

Job description

While our existing perception division handles state estimation and spatial mapping, this role is dedicated to one of the most critical bottlenecks in embodied AI: full-body kinematics, human locomotion, and human-scene interaction. We are hiring a Research Scientist to architect and train proprietary foundation models from scratch focused on 3D human body tracking and articulated pose estimation.

Your core mandate is twofold: building our in-house equivalents to cutting-edge 3D human body and mesh recovery architectures, and developing highly robust interaction models tailored for complex, real-world environments characterized by severe occlusions and dynamic motion. Beyond these core pillars, you will serve as a lead problem-solver for emergent perception challenges as our hardware and downstream robotics needs evolve.

To achieve this, we can provide a massive, continuous stream of high-quality, proprietary ground-truth human motion data captured by our infrastructure. You will use this data advantage to train networks that surpass current public baselines, owning the complete human-scene perception loop for our data engine.

What You’ll Work On

Architecting Proprietary Articulation Models

  • Zero-to-One Model Development: Design, implement, and train state-of-the-art networks for 3D human pose estimation, dense full-body mesh recovery, and kinematic tracking.
  • Large-Scale Distributed Training: Scale multi-view and temporal ML architectures across multi-GPU clusters to handle massive, multi-modal datasets of humans navigating and interacting with their environments.
  • Loss & Architecture Innovation: Push the boundaries of current paradigms by developing novel loss functions that enforce biomechanical constraints, temporal smoothness, postural balance, and physical plausibility.

Human-Scene Interaction (HSI) & Complex Motion Modeling

  • Dynamic Scene Understanding: Build and train custom architectures capable of handling extreme motion blur, severe self-occlusion, and multi-person crowding inherent in real-world human behavior.
  • Allocentric & Egocentric Tracking: Use your models to track human bodies through complex spaces, mapping foot-to-ground contact, joint torques, and environmental affordances to provide rich regularization for downstream action-conditioned robotics models (especially humanoid robots).

Emergent Perception R&D

  • Rapid Prototyping: Tackle novel, unmapped AI challenges as they arise. You will rapidly prototype and deploy new models for tasks spanning fine-grained action segmentation, intent prediction, and novel hardware sensor integrations.
  • Agile Problem Solving: Pivot to resolve sudden algorithmic bottlenecks in the data engine, adapting the latest research to unblock new product capabilities for our robotics customers., * The Data Advantage: You will have access to a scale and quality of proprietary spatial and temporal ground truth for human motion that most academic researchers only dream of.
  • Pure R&D & Model Ownership: You are not maintaining legacy systems; you are given a blank slate and the compute resources to build the state-of-the-art.
  • High Impact: The kinematic priors and interaction models you architect will directly define how the next generation of embodied AI agents-from mobile manipulators to humanoid robots-learn to physically navigate, balance, and interact with the world.

Requirements

  • Deep expertise in Deep Learning, 3D Computer Vision, and specifically Articulated Tracking / Human Body Pose Estimation.
  • Proven experience training large-scale vision models from scratch, not just running inference or fine-tuning existing checkpoints.
  • Strong theoretical and practical understanding of parametric human body models (e.g., SMPL, SMPL-X, GHUM, MHR, SOMA-X), inverse kinematics, and dense mesh estimation.
  • Mastery of PyTorch and deep learning scaling frameworks.
  • Experience handling and curating massive, multi-terabyte image and video datasets for training.
  • Comfortable operating in a fast-paced environment where priorities can shift rapidly to capitalize on new research or hardware capabilities.

Strong Signals:

  • First-author publications in top-tier venues (CVPR, ICCV, ECCV, NeurIPS) focusing on 3D human pose tracking, human-scene interaction (HSI), human motion capture, or human mesh recovery.
  • Specific experience working with massive human motion and interaction datasets (e.g., AMASS, Human3.6M, EgoBody, PROX) and solving the unique optimization challenges they present.

About the company

Mecka AI is building the data infrastructure layer for robotics and embodied AI. We design and operate global systems for data capture, data labeling, and hardware-enabled workflows used by leading AI labs and robotics companies to train and validate humanoid and embodied AI systems. We work closely with frontier robotics teams to bridge real-world data, simulation, learning-based systems, and deployed hardware., Mecka uses artificial intelligence (AI) responsibly to support administrative and efficiency-focused aspects of our recruitment process. This includes activities such as drafting job descriptions, generating interview questions, note-taking and recordings, and supporting sourcing and scheduling workflows. All candidate evaluations, interviews, and hiring decisions are made by members of the Mecka team. While AI tools may assist with screening and assessment, they do not replace human judgment in selection decisions. Our use of AI is intended to streamline routine tasks, improve consistency, and enhance the overall candidate experience. We are committed to upholding principles of fairness, transparency, and accountability in all hiring activities. Mecka regularly reviews its recruitment practices to mitigate bias and to ensure alignment with applicable laws and evolving best practices.

Compensation Range: $150K - $200K

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:47 min

Contrasting classical machine learning with deep learning approaches

Adrian Spataru +1 · LIVE

8:06 min

Core hardware and software components of robotics

Florian Gather Florian Gather +1 · World Congress 2026 Europe

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

4:29 min

Expanding recognition to human body pose tracking

Milan Todorovic · LIVE

2:54 min

Hallucinating valid trajectories using video diffusion models

Alexander Schwarz Alexander Schwarz · World Congress 2025

1:25 min

Distinguishing artificial intelligence from deep learning

Sam Witteveen · Coffee With Developers

Videos

See all

Related articles

See all