> Markdown version of [/jobs/ext/3376510-research-scientist-computer-vision-body-pose-detection](https://www.wearedevelopers.com/jobs/ext/3376510-research-scientist-computer-vision-body-pose-detection). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Research Scientist - Computer Vision (Body Pose Detection) - **Company:** MECKA ASSOC - **Location:** New York, NY, United States - **Salary:** $150,000.0 - $200,000.0 - **Contract:** Permanent contract - **Skills:** Agile Methodology, Artificial Intelligence, Computer Vision, Computer Clusters, Information Engineering, Distributed Computing Environment, Kinematics, Motion Capture, Rapid Prototyping Process, Pytorch, Deep Learning, Legacy Systems - **Published:** September 27, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=2268e1ef844392c6 ## About the Role * Deep expertise in Deep Learning, 3D Computer Vision, and specifically Articulated Tracking / Human Body Pose Estimation. * Proven experience training large-scale vision models from scratch, not just running inference or fine-tuning existing checkpoints. * Strong theoretical and practical understanding of parametric human body models (e.g., SMPL, SMPL-X, GHUM, MHR, SOMA-X), inverse kinematics, and dense mesh estimation. * Mastery of PyTorch and deep learning scaling frameworks. * Experience handling and curating massive, multi-terabyte image and video datasets for training. * Comfortable operating in a fast-paced environment where priorities can shift rapidly to capitalize on new research or hardware capabilities. Strong Signals: * First-author publications in top-tier venues (CVPR, ICCV, ECCV, NeurIPS) focusing on 3D human pose tracking, human-scene interaction (HSI), human motion capture, or human mesh recovery. * Specific experience working with massive human motion and interaction datasets (e.g., AMASS, Human3.6M, EgoBody, PROX) and solving the unique optimization challenges they present. ## Description While our existing perception division handles state estimation and spatial mapping, this role is dedicated to one of the most critical bottlenecks in embodied AI: full-body kinematics, human locomotion, and human-scene interaction. We are hiring a Research Scientist to architect and train proprietary foundation models from scratch focused on 3D human body tracking and articulated pose estimation. Your core mandate is twofold: building our in-house equivalents to cutting-edge 3D human body and mesh recovery architectures, and developing highly robust interaction models tailored for complex, real-world environments characterized by severe occlusions and dynamic motion. Beyond these core pillars, you will serve as a lead problem-solver for emergent perception challenges as our hardware and downstream robotics needs evolve. To achieve this, we can provide a massive, continuous stream of high-quality, proprietary ground-truth human motion data captured by our infrastructure. You will use this data advantage to train networks that surpass current public baselines, owning the complete human-scene perception loop for our data engine. What You'll Work On Architecting Proprietary Articulation Models * Zero-to-One Model Development: Design, implement, and train state-of-the-art networks for 3D human pose estimation, dense full-body mesh recovery, and kinematic tracking. * Large-Scale Distributed Training: Scale multi-view and temporal ML architectures across multi-GPU clusters to handle massive, multi-modal datasets of humans navigating and interacting with their environments. * Loss & Architecture Innovation: Push the boundaries of current paradigms by developing novel loss functions that enforce biomechanical constraints, temporal smoothness, postural balance, and physical plausibility. Human-Scene Interaction (HSI) & Complex Motion Modeling * Dynamic Scene Understanding: Build and train custom architectures capable of handling extreme motion blur, severe self-occlusion, and multi-person crowding inherent in real-world human behavior. * Allocentric & Egocentric Tracking: Use your models to track human bodies through complex spaces, mapping foot-to-ground contact, joint torques, and environmental affordances to provide rich regularization for downstream action-conditioned robotics models (especially humanoid robots). Emergent Perception R&D * Rapid Prototyping: Tackle novel, unmapped AI challenges as they arise. You will rapidly prototype and deploy new models for tasks spanning fine-grained action segmentation, intent prediction, and novel hardware sensor integrations. * Agile Problem Solving: Pivot to resolve sudden algorithmic bottlenecks in the data engine, adapting the latest research to unblock new product capabilities for our robotics customers., * The Data Advantage: You will have access to a scale and quality of proprietary spatial and temporal ground truth for human motion that most academic researchers only dream of. * Pure R&D & Model Ownership: You are not maintaining legacy systems; you are given a blank slate and the compute resources to build the state-of-the-art. * High Impact: The kinematic priors and interaction models you architect will directly define how the next generation of embodied AI agents-from mobile manipulators to humanoid robots-learn to physically navigate, balance, and interact with the world. ## Related Videos - [How I built my own intelligent Robot Arm from Scratch](https://www.wearedevelopers.com/videos/100097-how-i-built-my-own-intelligent-robot-arm-from-scratch) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [Detect Hand Pose with Vision](https://www.wearedevelopers.com/videos/135-detect-hand-pose-with-vision) - [The pitfalls of Deep Learning - When Neural Networks are not the solution](https://www.wearedevelopers.com/videos/14-the-pitfalls-of-deep-learning-when-neural-networks-are-not-the-solution) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Getting Started with Machine Learning](https://www.wearedevelopers.com/videos/260-getting-started-with-machine-learning) ## Related Articles - [What’s in The box? – Unboxing The DeepFace](https://www.wearedevelopers.com/magazine/117-what-s-in-the-box-unboxing-the-deepface) - [ I Gave a Video Editor More Autonomy Than a Trading Bot. On Purpose.](https://www.wearedevelopers.com/magazine/773-i-gave-a-video-editor-more-autonomy-than-a-trading-bot-on-purpose) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this)