Devops Engineer - Ai Model Evaluator
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
About the RoleMercor is partnering with a leading AI research lab to support a Frontier Code Agents project.Contributors help evaluate and improve frontier AI coding models through structured technical assessments.The work focuses on realistic infrastructure engineering workflows and model evaluation.Spots are limited and filling quickly on a first come, first serve basis.What You’ll DoUse frontier AI coding agents to complete and evaluate complex infrastructure engineering tasks.Review model-generated implementations involving cloud platforms, Kubernetes, CI/CD systems, observability, and infrastructure automation.Identify bugs, edge cases, reliability issues, and failure modes.Compare outputs from multiple frontier models and assess their strengths and weaknesses.Apply professional engineering judgment to realistic infrastructure engineering scenarios.Time CommitmentSprint based project that runs in *** hour stretches based on client requirement.Compensation$400 per accepted task.Typical tasks take approximately 2-3 hours after ramp-up.Compensation is tied to accepted work.Who Should Apply2+ years of professional DevOps, SRE, or Cloud Engineering experience.Experience with AWS, Azure, GCP, Kubernetes, Terraform, CI/CD pipelines, or observability tooling.Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools.Ability to evaluate model-generated infrastructure and reliability engineering solutions.xqbhyrxExperience supporting production-scale systems is preferred.#J-***-Ljbffr
Requirements
Who Should Apply2+ years of professional DevOps, SRE, or Cloud Engineering experience. Experience with AWS, Azure, GCP, Kubernetes, Terraform, CI/CD pipelines, or observability tooling. Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools. Ability to evaluate model-generated infrastructure and reliability engineering solutions. xqbhyrx Experience supporting production-scale systems is preferred. #J-*****-Ljbffr
About the company
About the RoleMercor is partnering with a leading AI research lab to support a Frontier Code Agents project. Contributors help evaluate and improve frontier AI coding models through structured technical assessments. The work focuses on realistic infrastructure engineering workflows and model evaluation. Spots are limited and filling quickly on a first come, first serve basis. What You’ll DoUse frontier AI coding agents to complete and evaluate complex infrastructure engineering tasks. Review model-generated implementations involving cloud platforms, Kubernetes, CI/CD systems, observability, and infrastructure automation. Identify bugs, edge cases, reliability issues, and failure modes. Compare outputs from multiple frontier models and assess their strengths and weaknesses. Apply professional engineering judgment to realistic infrastructure engineering scenarios. Time CommitmentSprint based project that runs in ***** hour stretches based on client requirement. Compensation$400 per accepted task. Typical tasks take approximately 2-3 hours after ramp-up.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.buscojobs.com.esGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Dev Digest 121 - AI goes offline
Dev Digest 137 - AI'm not sure about this
Navigating the AI Shift
MLOps And AI Driven Development