MLOps Engineer, LLM Systems

LAKE ST LLC
Lake Saint Croix Beach, MN, United States
4 days ago
Apply on www.juju.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Compensation
$187,200.0 - $249,600.0
Working hours
Regular working hours
Job source

Tech stack

Training Data Profiling Nvidia CUDA Software Debugging Linux Kernel Pytorch Large Language Models Low Latency Machine Learning Operations TensorRT

Job description

  • Design challenging, domain-relevant MLOps and ML systems tasks in GPU kernels, profiling, debugging, and inference serving, then produce accurate, well-structured solutions.
  • Evaluate technical tasks and solutions, providing clear written feedback that can withstand detailed review.
  • Support research and engineering teams in closing knowledge gaps and improving model performance across ML systems, training infrastructure, and framework-level subjects.
  • Create detailed guidelines and evaluation rubrics for kernel optimization, profiler-output interpretation, distributed-systems reasoning, and serving throughput and latency trade-offs.
  • Partner with subject matter experts to maintain consistent, accurate training data.

Requirements

  • At least 2 years of hands-on professional experience in ML systems, ML infrastructure, model serving, or GPU and accelerator performance engineering.
  • Experience in at least one of the following areas, with experience across multiple areas strongly preferred: custom GPU kernel development or optimization using CUDA, Triton, or Pallas; profiling and trace analysis using Kineto, torch.profiler, Nsight, XLA, or JAX profiler; debugging distributed or accelerator-bound workloads; or serving LLMs at scale using vLLM, SGLang, TensorRT-LLM, Ray Serve, KV cache, paged attention, or continuous batching.
  • Production experience with JAX and/or PyTorch. Framework-level expertise in custom operators, distributed training with FSDP, DDP, DeepSpeed, or Megatron, or compiler and graph-level work is preferred.
  • Familiarity with A100, H100, B200, or TPU accelerators, including the ability to assess throughput, latency, and memory trade-offs.
  • Demonstrable career progression, strong written communication, and the ability to explain complex technical decisions clearly.
  • This is a systems-focused position, not an applied modeling or data science role.

Benefits & conditions

Work Terms

  • Hourly W-2 employment with placement on an extended workforce team supporting a leading AI lab.
  • Full-time, 40-hour-per-week weekday commitment. Candidates must be able to work without conflicting engagements or other concurrent work commitments.
  • Available to candidates located in Canada, the United Kingdom, or the United States.

Compensation

$90 to $120 per hour.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.juju.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

5:28 min

Defining MLOps and its role in production systems

Hauke Brammer · World Congress 2023

47 sec

Profiling native execution calls with async-profiler

Gonzalo Ortiz Jaureguizar Gonzalo Ortiz Jaureguizar · World Congress 2026 Europe

4:52 min

Essential phases in building and refining language models

Anshul Jindal Anshul Jindal +1 · World Congress 2025

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

2:44 min

Defining core roles and responsibilities in MLOps teams

Bas Geerdink · LIVE

2:17 min

Comparing code profiling with surface level monitoring

Jérôme Vieilledent · LIVE

Videos

See all

Related articles

See all