Principal Machine Learning Engineer

On behalf of Next Deavor
New York, NY, United States
about 1 month ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Compensation
$200,000.0 - $250,000.0
Working hours
Regular working hours
Job source

Tech stack

Application Release Automation Cloud Computing Continuous Integration Distributed Computing Environment Python (Programming Language) Machine Learning Language Modeling Software Engineering Management of Software Versions Pytorch Large Language Models Machine Learning Operations
+1 more
TensorRT

Job description

You will own the ML infrastructure that turns research into reliable, real-time compliance enforcement systems, driving model training, evaluation, and production serving. You will partner closely with research stakeholders and engineering peers to ship reproducible pipelines and low-latency serving; the role is Hybrid (3 days onsite) in the New York City Metro area.

Here’s How You’ll Make an Impact on the Team

  • Build and own training pipelines: data preparation, reproducible fine-tuning runs, experiment tracking, and release automation
  • Build evaluation infrastructure: automated eval runs, regression gates, dashboards, and dataset versioning
  • Own model serving in production: low-latency inference, batching, optimization, autoscaling, and cost management
  • Ship model updates safely with versioning, canarying, rollback, and drift monitoring
  • Create repeatable workflows to adapt models to new domains and customer needs
  • Turn expert labels and reviewer feedback into clean training and evaluation data
  • Set the engineering bar for ML infrastructure as the team grows

Requirements

  • 8+ years of software engineering experience, including 4+ years building infrastructure for ML or LLM systems in production
  • Hands-on experience with the modern LLM stack: PyTorch, distributed training, fine-tuning at scale (e.g., LoRA, SFT), and inference engines such as vLLM or TensorRT-LLM
  • Experience building eval harnesses, regression gates, or dataset pipelines; strong understanding of precision, recall, and calibration
  • Proven ownership of production model serving with real latency, reliability, and cost constraints
  • Strong fundamentals in Python, containers, CI/CD, cloud infrastructure, and observability
  • Ability to scope work, ship frequently, and make pragmatic build-vs-buy decisions
  • Experience collaborating tightly with research partners and defining clear interfaces

Here’s What Else Might Help You Out

  • Experience productionizing small or specialized language models
  • Experience with structured-output serving or constrained decoding in production
  • Prior work in regulated or high-stakes domains (fintech, healthcare, legal, trust and safety)
  • Experience deploying models into customer-controlled environments

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Open-source community and machine learning frameworks

Gian Marco Iodice Gian Marco Iodice · World Congress 2025

2:27 min

Introduction to WebAssembly in a cloud computing context

Edo Edo · World Congress 2024

4:52 min

Essential phases in building and refining language models

Anshul Jindal Anshul Jindal +1 · World Congress 2025

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

2:15 min

Bridging the gap between model management and devops

Joy Joy · World Congress 2024

2:51 min

Alibaba Cloud developer resources and cloud computing training

Cheng Zhang · LIVE

Videos

See all

Related articles

See all