Senior Machine Learning Engineer

Placement, Ltd
United States
21 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
2 years minimum
Compensation
$160,000.0 - $260,000.0
Working hours
Regular working hours
Job source

Tech stack

Data Deduplication Fault Tolerance Machine Learning Node.Js Software Deployment Large Language Models Parallel Computation TensorRT Data Pipelines

Job description

We’re looking for a Machine Learning Engineer who can own the full lifecycle of alarge language model in production - from training and fine-tuning throughdeployment at scale and rigorous evaluation of what actually ships. This isn’t aresearch-only role and it isn’t a pure infrastructure role: you’ll need to begenuinely fluent in all three, because the hardest problems in this space live atthe seams between them - a training decision that breaks serving latency, adeployment optimization that silently degrades output quality, an eval result thatdoesn’t predict real-world behavior.

What you’ll do:

Training & fine-tuning

    • Design and execute pretraining, continued pretraining, and fine-tuning runs (SFT, DPO/RLHF-style alignment, LoRA/QLoRA and full-parameter approaches) against clear, measurable objectives
    • Own data pipeline decisions that materially affect model quality - curation, deduplication, mixture weighting, and contamination checks against eval sets- Run and interpret distributed training (multi-GPU, multi-node) using frameworks such as FSDP, DeepSpeed, or Megatron-style parallelism, and diagnose failures that only show up at scale (loss spikes, stragglers, checkpoint corruption)
    • Make and defend real tradeoffs between model size, training cost, and downstream performance

Deployment

  • Take a trained model to production: quantization, batching strategy, KV-cache management, and serving framework selection (e.g. vLLM, TensorRT-LLM, TGI) with explicit latency/throughput/cost targets, not just “make it run”
  • Design for the failure modes specific to LLM serving - tail latency under load, graceful degradation, prompt injection surface area, and safe fallback behavior
  • Build the operational muscle around this: monitoring, alerting, and rollback paths for a model in production, treated with the same rigor as any other critical service, not as a one-off notebook export

Assessment & evaluation

  • Build and maintain evaluation harnesses that go beyond running published benchmarks - including task-specific eval sets that reflect your actual product’s use cases, not just leaderboard performance
  • Design human evaluation protocols where automated metrics fall short, and know which is which
  • Own regression detection: catching quality drops introduced by a new checkpoint, a prompt template change, or a serving optimization before they reach users- Contribute to safety and robustness evaluation - hallucination rate, adversarial/ red-team testing, and behavior under distribution shift - as a first-class part of the release process, not an afterthought

Requirements

  • 4+ years of applied ML engineering experience, with at least **2 years working directly on large language models in a production context
  • Real production deployment experience - you’ve shipped a model that served live traffic, and you can talk about the latency/cost/quality tradeoffs you made- Strong software engineering fundamentals
  • Fluency with the modern LLM tooling landscape (training frameworks, serving frameworks, eval tooling)
  • Comfort with ambiguity - you’ll be asked to define what “good” means for a model behavior that doesn’t have an established benchmark

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:20 min

Combating human workforce shortages with specialized language models

Markus Hacker Markus Hacker +3 ¡ World Congress 2024

6:08 min

Applying software engineering environments and testing to data pipelines

Matthias Niehoff Matthias Niehoff ¡ World Congress 2024

4:52 min

Essential phases in building and refining language models

Anshul Jindal Anshul Jindal +1 ¡ World Congress 2025

45 sec

Working securely with Node.js path application programming interfaces

Sonya Moisset ¡ World Congress 2023

2:08 min

Applying large language models to infrastructure tasks

Alfonso Sandoval Rosas Alfonso Sandoval Rosas ¡ Europe 2026 Virtual

47 sec

Building modern data pipelines for legacy exports

Dr. Alexander Wachtel Dr. Alexander Wachtel +1 ¡ World Congress 2025

Videos

See all

Related articles

See all