> Markdown version of [/jobs/ext/1402929-senior-machine-learning-engineer](https://www.wearedevelopers.com/jobs/ext/1402929-senior-machine-learning-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Machine Learning Engineer - **Company:** Placement, Ltd - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $160,000.0 - $260,000.0 - **Contract:** Permanent contract - **Skills:** Data Deduplication, Fault Tolerance, Machine Learning, Node.Js, Software Deployment, Large Language Models, Parallel Computation, TensorRT, Data Pipelines - **Published:** July 23, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=a299dca3b052e682 ## About the Role * 4+ years of applied ML engineering experience, with at least **2 years working directly on large language models in a production context * Real production deployment experience - you've shipped a model that served live traffic, and you can talk about the latency/cost/quality tradeoffs you made- Strong software engineering fundamentals * Fluency with the modern LLM tooling landscape (training frameworks, serving frameworks, eval tooling) * Comfort with ambiguity - you'll be asked to define what "good" means for a model behavior that doesn't have an established benchmark ## Description We're looking for a Machine Learning Engineer who can own the full lifecycle of alarge language model in production - from training and fine-tuning throughdeployment at scale and rigorous evaluation of what actually ships. This isn't aresearch-only role and it isn't a pure infrastructure role: you'll need to begenuinely fluent in all three, because the hardest problems in this space live atthe seams between them - a training decision that breaks serving latency, adeployment optimization that silently degrades output quality, an eval result thatdoesn't predict real-world behavior. What you'll do: Training & fine-tuning * - Design and execute pretraining, continued pretraining, and fine-tuning runs (SFT, DPO/RLHF-style alignment, LoRA/QLoRA and full-parameter approaches) against clear, measurable objectives * - Own data pipeline decisions that materially affect model quality - curation, deduplication, mixture weighting, and contamination checks against eval sets- Run and interpret distributed training (multi-GPU, multi-node) using frameworks such as FSDP, DeepSpeed, or Megatron-style parallelism, and diagnose failures that only show up at scale (loss spikes, stragglers, checkpoint corruption) * - Make and defend real tradeoffs between model size, training cost, and downstream performance Deployment * Take a trained model to production: quantization, batching strategy, KV-cache management, and serving framework selection (e.g. vLLM, TensorRT-LLM, TGI) with explicit latency/throughput/cost targets, not just "make it run" * Design for the failure modes specific to LLM serving - tail latency under load, graceful degradation, prompt injection surface area, and safe fallback behavior * Build the operational muscle around this: monitoring, alerting, and rollback paths for a model in production, treated with the same rigor as any other critical service, not as a one-off notebook export Assessment & evaluation * Build and maintain evaluation harnesses that go beyond running published benchmarks - including task-specific eval sets that reflect your actual product's use cases, not just leaderboard performance * Design human evaluation protocols where automated metrics fall short, and know which is which * Own regression detection: catching quality drops introduced by a new checkpoint, a prompt template change, or a serving optimization before they reach users- Contribute to safety and robustness evaluation - hallucination rate, adversarial/ red-team testing, and behavior under distribution shift - as a first-class part of the release process, not an afterthought ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Why and when should we consider Stream Processing frameworks in our solutions](https://www.wearedevelopers.com/videos/1085-why-and-when-should-we-consider-stream-processing-frameworks-in-our-solutions) - [Stop using Node.js like in 2020! What changed and what you can do today with Node.js](https://www.wearedevelopers.com/videos/100011-stop-using-node-js-like-in-2020-what-changed-and-what-you-can-do-today-with-node-js) - [DevOps for Machine Learning](https://www.wearedevelopers.com/videos/179-devops-for-machine-learning) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Python-Based Data Streaming Pipelines Within Minutes](https://www.wearedevelopers.com/videos/1233-python-based-data-streaming-pipelines-within-minutes) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [The Best Large Language Models on The Market](https://www.wearedevelopers.com/magazine/319-the-best-large-language-models-on-the-market) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)