> Markdown version of [/jobs/ext/2627840-principal-machine-learning-engineer](https://www.wearedevelopers.com/jobs/ext/2627840-principal-machine-learning-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Machine Learning Engineer - **Company:** On behalf of Next Deavor - **Location:** New York, NY, United States - **Experience:** Experienced - **Salary:** $200,000.0 - $250,000.0 - **Contract:** Permanent contract - **Skills:** Application Release Automation, Cloud Computing, Continuous Integration, Distributed Computing Environment, Python (Programming Language), Machine Learning, Language Modeling, Software Engineering, Management of Software Versions, Pytorch, Large Language Models, Machine Learning Operations, TensorRT - **Published:** August 4, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=f346cad200169df2 ## About the Role * 8+ years of software engineering experience, including 4+ years building infrastructure for ML or LLM systems in production * Hands-on experience with the modern LLM stack: PyTorch, distributed training, fine-tuning at scale (e.g., LoRA, SFT), and inference engines such as vLLM or TensorRT-LLM * Experience building eval harnesses, regression gates, or dataset pipelines; strong understanding of precision, recall, and calibration * Proven ownership of production model serving with real latency, reliability, and cost constraints * Strong fundamentals in Python, containers, CI/CD, cloud infrastructure, and observability * Ability to scope work, ship frequently, and make pragmatic build-vs-buy decisions * Experience collaborating tightly with research partners and defining clear interfaces Here's What Else Might Help You Out * Experience productionizing small or specialized language models * Experience with structured-output serving or constrained decoding in production * Prior work in regulated or high-stakes domains (fintech, healthcare, legal, trust and safety) * Experience deploying models into customer-controlled environments ## Description You will own the ML infrastructure that turns research into reliable, real-time compliance enforcement systems, driving model training, evaluation, and production serving. You will partner closely with research stakeholders and engineering peers to ship reproducible pipelines and low-latency serving; the role is Hybrid (3 days onsite) in the New York City Metro area. Here's How You'll Make an Impact on the Team * Build and own training pipelines: data preparation, reproducible fine-tuning runs, experiment tracking, and release automation * Build evaluation infrastructure: automated eval runs, regression gates, dashboards, and dataset versioning * Own model serving in production: low-latency inference, batching, optimization, autoscaling, and cost management * Ship model updates safely with versioning, canarying, rollback, and drift monitoring * Create repeatable workflows to adapt models to new domains and customer needs * Turn expert labels and reviewer feedback into clean training and evaluation data * Set the engineering bar for ML infrastructure as the team grows ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [DevOps for Machine Learning](https://www.wearedevelopers.com/videos/179-devops-for-machine-learning) - [Green Cloud Computing](https://www.wearedevelopers.com/videos/592-green-cloud-computing) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path)