Senior Machine Learning Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
Weâre looking for a Machine Learning Engineer who can own the full lifecycle of alarge language model in production - from training and fine-tuning throughdeployment at scale and rigorous evaluation of what actually ships. This isnât aresearch-only role and it isnât a pure infrastructure role: youâll need to begenuinely fluent in all three, because the hardest problems in this space live atthe seams between them - a training decision that breaks serving latency, adeployment optimization that silently degrades output quality, an eval result thatdoesnât predict real-world behavior.
What youâll do:
Training & fine-tuning
-
- Design and execute pretraining, continued pretraining, and fine-tuning runs (SFT, DPO/RLHF-style alignment, LoRA/QLoRA and full-parameter approaches) against clear, measurable objectives
-
- Own data pipeline decisions that materially affect model quality - curation, deduplication, mixture weighting, and contamination checks against eval sets- Run and interpret distributed training (multi-GPU, multi-node) using frameworks such as FSDP, DeepSpeed, or Megatron-style parallelism, and diagnose failures that only show up at scale (loss spikes, stragglers, checkpoint corruption)
-
- Make and defend real tradeoffs between model size, training cost, and downstream performance
Deployment
- Take a trained model to production: quantization, batching strategy, KV-cache management, and serving framework selection (e.g. vLLM, TensorRT-LLM, TGI) with explicit latency/throughput/cost targets, not just âmake it runâ
- Design for the failure modes specific to LLM serving - tail latency under load, graceful degradation, prompt injection surface area, and safe fallback behavior
- Build the operational muscle around this: monitoring, alerting, and rollback paths for a model in production, treated with the same rigor as any other critical service, not as a one-off notebook export
Assessment & evaluation
- Build and maintain evaluation harnesses that go beyond running published benchmarks - including task-specific eval sets that reflect your actual productâs use cases, not just leaderboard performance
- Design human evaluation protocols where automated metrics fall short, and know which is which
- Own regression detection: catching quality drops introduced by a new checkpoint, a prompt template change, or a serving optimization before they reach users- Contribute to safety and robustness evaluation - hallucination rate, adversarial/ red-team testing, and behavior under distribution shift - as a first-class part of the release process, not an afterthought
Requirements
- 4+ years of applied ML engineering experience, with at least **2 years working directly on large language models in a production context
- Real production deployment experience - youâve shipped a model that served live traffic, and you can talk about the latency/cost/quality tradeoffs you made- Strong software engineering fundamentals
- Fluency with the modern LLM tooling landscape (training frameworks, serving frameworks, eval tooling)
- Comfort with ambiguity - youâll be asked to define what âgoodâ means for a model behavior that doesnât have an established benchmark
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.indeed.comGood distractions
Talks and stories from around this role â technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
MLops â Deploying, Maintaining And Evolving Machine Learning Models in Production
The Best Large Language Models on The Market
How to Become an AI Engineer
MLOps â Whatâs the deal behind it?