> Markdown version of [/jobs/ext/2713674-machine-learning-engineer-frontier-ai-evaluation-contract](https://www.wearedevelopers.com/jobs/ext/2713674-machine-learning-engineer-frontier-ai-evaluation-contract). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Engineer, Frontier AI Evaluation (Contract) - **Company:** Cobalt LLP - **Location:** Piedmont, CA, United States - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Software Debugging, Python (Programming Language), Machine Learning, Pytorch, Deep Learning, Machine Learning Operations, Data Pipelines - **Published:** September 4, 2026 - **Apply:** https://www.disabledperson.com/jobs/74813014-machine-learning-engineer-frontier-ai-evaluation-contract ## About the Role This opportunity is suited to practitioners rather than only researchers: ML engineers, applied scientists, MLOps and platform engineers, and data engineers who have trained, deployed, and maintained models in production. A PhD is welcome but not required, and hands-on delivery experience counts for more here than publication record. You do not need prior experience in data annotation or AI research. What matters is that you can diagnose why a pipeline or a training run is failing, decide what the right fix is, and explain both clearly enough for another engineer to follow., * Several years of hands-on experience building, training, and deploying machine learning systems in production, with a track record you can point to * Strong coding ability in Python, plus working command of at least one deep learning framework such as PyTorch or JAX, and comfort reading unfamiliar codebases * Depth in at least one area, for example large-scale training and distributed compute, data pipelines and feature infrastructure, model serving and inference optimization, evaluation and monitoring, or fine-tuning and post-training workflows * Solid debugging discipline, including the ability to isolate a failure across data, model, and infrastructure rather than guessing at it * Ability to explain each step of your reasoning clearly in writing, and to produce work another engineer could reproduce and review ## Description Cobalt is seeking machine learning engineers to produce the expert reasoning, task environments, and evaluation data used to train and assess frontier AI models on real ML engineering work., * Produce written reasoning traces on real ML engineering tasks, capturing how you diagnose a failing training run, a data pipeline defect, or a serving regression, including what you rule out and why * Author non-trivial ML engineering problems and task environments with checks that verify success automatically, including multi-file and multi-step tasks * Evaluate model-generated ML code and configurations, ranking solutions, explaining what makes the stronger one stronger, and identifying the point at which the approach goes wrong * Assess whether a proposed solution actually addresses the failure, and identify fixes that pass the immediate check but mask the underlying problem, degrade performance, or would not survive review * Design rubrics and partial-credit criteria for scoring multistep engineering tasks, and classify observed failures into a consistent taxonomy Projects follow their own guidelines, formatting conventions, and quality standards, and you will work with feedback from reviewers and lab research teams. ## Related Videos - [Getting Started with Machine Learning](https://www.wearedevelopers.com/videos/260-getting-started-with-machine-learning) - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [Why and when should we consider Stream Processing frameworks in our solutions](https://www.wearedevelopers.com/videos/1085-why-and-when-should-we-consider-stream-processing-frameworks-in-our-solutions) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [Navigating the AI Revolution in Software Development](https://www.wearedevelopers.com/videos/1266-navigating-the-ai-revolution-in-software-development) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path)