Machine Learning Engineer

Theory Llc
Washington, DC, United States
13 days ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Automated Storage and Retrieval Systems Python (Programming Language) Machine Learning Pytorch Large Language Models HuggingFace Machine Learning Operations

Requirements

  • Three or more years building machine learning systems that other people relied on
  • Fluent Python and the modern ML stack: PyTorch or JAX, Hugging Face, experiment tracking
  • Practical experience evaluating LLMs beyond a single aggregate score
  • Able to explain a technical trade-off to a client who is not an engineer

Nice to have

  • Worked under FedRAMP, HIPAA, or an equivalent compliance regime
  • Published or presented on evaluation methodology
  • Experience with retrieval systems on private corpora

Powered by JazzHR

Theory AI

Benefits & conditions

  • Fully paid medical, dental and vision
  • 401(k) with employer contribution
  • Unlimited paid time off with a four-week minimum
  • $3,000 a year for conferences, courses and certifications
  • Home office budget and a choice of hardware
  • This is a Remote (work from home) position.

Most of the machine learning work that reaches production is not modelling. It is knowing which failures matter, building the evidence that shows whether they still happen, and being able to defend that evidence to someone whose job is to find the hole in it. You will own model development and evaluation on client engagements, largely in regulated settings: federal health, prime contractors, enterprises with data that cannot leave their boundary. That means working against real constraints on residency and access rather than a clean benchmark, and writing up what you found in a form a review board can read. This role suits someone who has shipped a model that other people depended on, and who has opinions about why aggregate accuracy is a poor summary of anything. What you will do

  • Design and run evaluations for client models, including refusal behaviour and the failure modes a review board will test
  • Build reproducible training and fine-tuning pipelines against regulated data
  • Turn ambiguous client questions into measurable criteria before any modelling starts
  • Write up methodology and findings for technical and non-technical audiences
  • Review the ground-truth work our annotation teams produce, and feed problems back into the guidelines

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:18 min

Exploring the tiered architecture of modern machine learning stacks

Kris Howard · LIVE

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

2:36 min

Applying supervised machine learning for practical rule extraction

Katja Träumner

7:10 min

Exploring pathways into the machine learning engineering field

Jose Luis Latorre Millas · LIVE

1:06 min

Compiling PyTorch environments for advanced time forecasting

Christoph Lohrmann Christoph Lohrmann +1 · World Congress 2026 Europe

1:57 min

Evolution of machine learning algorithms and computing hardware

Alexandra Waldherr · LIVE

Videos

See all

Related articles

See all