Principal Machine Learning Engineer
On behalf of Next Deavor
New York, NY, United States
about 1 month ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.indeed.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Compensation
$200,000.0 - $250,000.0
Working hours
Regular working hours
Job source
Tech stack
Application Release Automation
Cloud Computing
Continuous Integration
Distributed Computing Environment
Python (Programming Language)
Machine Learning
Language Modeling
Software Engineering
Management of Software Versions
Pytorch
Large Language Models
Machine Learning Operations
+1 more
TensorRT
Job description
You will own the ML infrastructure that turns research into reliable, real-time compliance enforcement systems, driving model training, evaluation, and production serving. You will partner closely with research stakeholders and engineering peers to ship reproducible pipelines and low-latency serving; the role is Hybrid (3 days onsite) in the New York City Metro area.
Here’s How You’ll Make an Impact on the Team
- Build and own training pipelines: data preparation, reproducible fine-tuning runs, experiment tracking, and release automation
- Build evaluation infrastructure: automated eval runs, regression gates, dashboards, and dataset versioning
- Own model serving in production: low-latency inference, batching, optimization, autoscaling, and cost management
- Ship model updates safely with versioning, canarying, rollback, and drift monitoring
- Create repeatable workflows to adapt models to new domains and customer needs
- Turn expert labels and reviewer feedback into clean training and evaluation data
- Set the engineering bar for ML infrastructure as the team grows
Requirements
- 8+ years of software engineering experience, including 4+ years building infrastructure for ML or LLM systems in production
- Hands-on experience with the modern LLM stack: PyTorch, distributed training, fine-tuning at scale (e.g., LoRA, SFT), and inference engines such as vLLM or TensorRT-LLM
- Experience building eval harnesses, regression gates, or dataset pipelines; strong understanding of precision, recall, and calibration
- Proven ownership of production model serving with real latency, reliability, and cost constraints
- Strong fundamentals in Python, containers, CI/CD, cloud infrastructure, and observability
- Ability to scope work, ship frequently, and make pragmatic build-vs-buy decisions
- Experience collaborating tightly with research partners and defining clear interfaces
Here’s What Else Might Help You Out
- Experience productionizing small or specialized language models
- Experience with structured-output serving or constrained decoding in production
- Prior work in regulated or high-stakes domains (fintech, healthcare, legal, trust and safety)
- Experience deploying models into customer-controlled environments
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.indeed.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
LM
Luis Minvielle
almost 3 years ago
BB
Benedikt Bischof
MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production
about 4 years ago
BB
Benedikt Bischof
MLOps – What’s the deal behind it?
almost 4 years ago
BB
Benedikt Bischof
MLOps And AI Driven Development
over 4 years ago
LM
Luis Minvielle
How to Become an AI Engineer
almost 3 years ago
EF
Elizabeth Fuentes Leone, AWS Developer Advocate, GenAI
From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path
10 months ago