Machine Learning Engineer, Frontier AI Evaluation (Contract)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
Cobalt is seeking machine learning engineers to produce the expert reasoning, task environments, and evaluation data used to train and assess frontier AI models on real ML engineering work., * Produce written reasoning traces on real ML engineering tasks, capturing how you diagnose a failing training run, a data pipeline defect, or a serving regression, including what you rule out and why
- Author non-trivial ML engineering problems and task environments with checks that verify success automatically, including multi-file and multi-step tasks
- Evaluate model-generated ML code and configurations, ranking solutions, explaining what makes the stronger one stronger, and identifying the point at which the approach goes wrong
- Assess whether a proposed solution actually addresses the failure, and identify fixes that pass the immediate check but mask the underlying problem, degrade performance, or would not survive review
- Design rubrics and partial-credit criteria for scoring multistep engineering tasks, and classify observed failures into a consistent taxonomy
Projects follow their own guidelines, formatting conventions, and quality standards, and you will work with feedback from reviewers and lab research teams.
Requirements
This opportunity is suited to practitioners rather than only researchers: ML engineers, applied scientists, MLOps and platform engineers, and data engineers who have trained, deployed, and maintained models in production. A PhD is welcome but not required, and hands-on delivery experience counts for more here than publication record.
You do not need prior experience in data annotation or AI research. What matters is that you can diagnose why a pipeline or a training run is failing, decide what the right fix is, and explain both clearly enough for another engineer to follow., * Several years of hands-on experience building, training, and deploying machine learning systems in production, with a track record you can point to
- Strong coding ability in Python, plus working command of at least one deep learning framework such as PyTorch or JAX, and comfort reading unfamiliar codebases
- Depth in at least one area, for example large-scale training and distributed compute, data pipelines and feature infrastructure, model serving and inference optimization, evaluation and monitoring, or fine-tuning and post-training workflows
- Solid debugging discipline, including the ability to isolate a failure across data, model, and infrastructure rather than guessing at it
- Ability to explain each step of your reasoning clearly in writing, and to produce work another engineer could reproduce and review
Benefits & conditions
- Grow professionally. Expand your influence through evaluation projects, advisory roles, and research collaborations, while developing a working understanding of how frontier models are trained and assessed.
- Work with a top-tier network. Collaborate with researchers and engineers from leading institutions and labs on high-impact, flexible work.
- Set your own schedule. Flexible 10 to 40 hour weeks that fit around your existing work and your life.
- Competitive pay. Rates vary by project and are determined by a number of factors, including scope, skillset, and experience.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
MLOps And AI Driven Development
MLOps – What’s the deal behind it?
What Are Large Language Models?
MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production