Senior Machine Learning Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
Weâre looking for a Machine Learning Engineer who can own the full lifecycle of alarge language model in production - from training and fine-tuning throughdeployment at scale and rigorous evaluation of what actually ships. This isnât aresearch-only role and it isnât a pure infrastructure role: youâll need to begenuinely fluent in all three, because the hardest problems in this space live atthe seams between them - a training decision that breaks serving latency, adeployment optimization that silently degrades output quality, an eval result thatdoesnât predict real-world behavior.
What youâll do:
Training & fine-tuning
-
- Design and execute pretraining, continued pretraining, and fine-tuning runs (SFT, DPO/RLHF-style alignment, LoRA/QLoRA and full-parameter approaches) against clear, measurable objectives
-
- Own data pipeline decisions that materially affect model quality - curation, deduplication, mixture weighting, and contamination checks against eval sets- Run and interpret distributed training (multi-GPU, multi-node) using frameworks such as FSDP, DeepSpeed, or Megatron-style parallelism, and diagnose failures that only show up at scale (loss spikes, stragglers, checkpoint corruption)
-
- Make and defend real tradeoffs between model size, training cost, and downstream performance
Deployment
- Take a trained model to production: quantization, batching strategy, KV-cache management, and serving framework selection (e.g. vLLM, TensorRT-LLM, TGI) with explicit latency/throughput/cost targets, not just âmake it runâ
- Design for the failure modes specific to LLM serving - tail latency under load, graceful degradation, prompt injection surface area, and safe fallback behavior
- Build the operational muscle around this: monitoring, alerting, and rollback paths for a model in production, treated with the same rigor as any other critical service, not as a one-off notebook export
Assessment & evaluation
- Build and maintain evaluation harnesses that go beyond running published benchmarks - including task-specific eval sets that reflect your actual productâs use cases, not just leaderboard performance
- Design human evaluation protocols where automated metrics fall short, and know which is which
- Own regression detection: catching quality drops introduced by a new checkpoint, a prompt template change, or a serving optimization before they reach users- Contribute to safety and robustness evaluation - hallucination rate, adversarial/ red-team testing, and behavior under distribution shift - as a first-class part of the release process, not an afterthought
Requirements
- 4+ years of applied ML engineering experience, with at least **2 years working directly on large language models in a production context
- Real production deployment experience - youâve shipped a model that served live traffic, and you can talk about the latency/cost/quality tradeoffs you made- Strong software engineering fundamentals
- Fluency with the modern LLM tooling landscape (training frameworks, serving frameworks, eval tooling)
- Comfort with ambiguity - youâll be asked to define what âgoodâ means for a model behavior that doesnât have an established benchmark
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role â technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
MLops â Deploying, Maintaining And Evolving Machine Learning Models in Production
The Best Large Language Models on The Market
How to Become an AI Engineer
MLOps â Whatâs the deal behind it?