MLOps Engineer, LLM Systems
LAKE ST LLC
Lake Saint Croix Beach, MN, United States
4 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.juju.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Compensation
$187,200.0 - $249,600.0
Working hours
Regular working hours
Job source
Tech stack
Training Data
Profiling
Nvidia CUDA
Software Debugging
Linux Kernel
Pytorch
Large Language Models
Low Latency
Machine Learning Operations
TensorRT
Job description
- Design challenging, domain-relevant MLOps and ML systems tasks in GPU kernels, profiling, debugging, and inference serving, then produce accurate, well-structured solutions.
- Evaluate technical tasks and solutions, providing clear written feedback that can withstand detailed review.
- Support research and engineering teams in closing knowledge gaps and improving model performance across ML systems, training infrastructure, and framework-level subjects.
- Create detailed guidelines and evaluation rubrics for kernel optimization, profiler-output interpretation, distributed-systems reasoning, and serving throughput and latency trade-offs.
- Partner with subject matter experts to maintain consistent, accurate training data.
Requirements
- At least 2 years of hands-on professional experience in ML systems, ML infrastructure, model serving, or GPU and accelerator performance engineering.
- Experience in at least one of the following areas, with experience across multiple areas strongly preferred: custom GPU kernel development or optimization using CUDA, Triton, or Pallas; profiling and trace analysis using Kineto, torch.profiler, Nsight, XLA, or JAX profiler; debugging distributed or accelerator-bound workloads; or serving LLMs at scale using vLLM, SGLang, TensorRT-LLM, Ray Serve, KV cache, paged attention, or continuous batching.
- Production experience with JAX and/or PyTorch. Framework-level expertise in custom operators, distributed training with FSDP, DDP, DeepSpeed, or Megatron, or compiler and graph-level work is preferred.
- Familiarity with A100, H100, B200, or TPU accelerators, including the ability to assess throughput, latency, and memory trade-offs.
- Demonstrable career progression, strong written communication, and the ability to explain complex technical decisions clearly.
- This is a systems-focused position, not an applied modeling or data science role.
Benefits & conditions
Work Terms
- Hourly W-2 employment with placement on an extended workforce team supporting a leading AI lab.
- Full-time, 40-hour-per-week weekday commitment. Candidates must be able to work without conflicting engagements or other concurrent work commitments.
- Available to candidates located in Canada, the United Kingdom, or the United States.
Compensation
$90 to $120 per hour.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.juju.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
LM
Luis Minvielle
almost 3 years ago
BB
Benedikt Bischof
MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production
about 4 years ago
BB
Benedikt Bischof
MLOps – What’s the deal behind it?
almost 4 years ago
BB
Benedikt Bischof
MLOps And AI Driven Development
over 4 years ago
KD
Krissy Davis
The Best Large Language Models on The Market
almost 3 years ago
EM
Eli McGarvie
Highest Paying Tech Companies for Developers
over 3 years ago