Enterprise Agent Development Platform project
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Requirements
behavioral checks Define evaluation criteria, metrics, thresholds, and acceptance rules Evaluate agent behavior across individual responses, tool calls, and complete workflows Work with OpenTelemetry traces and spans as evaluation data Integrate evaluations into CI/CD pipelines and automated deployment gates Enable continuous quality monitoring of solutions in production Establish reusable evaluation patterns and engineering standards Work closely with AI Engineers, Architects, and Platform Engineers to embed quality into the development process Requirements 5+ years of experience in ML Engineering, AI Engineering, or AI Platform Engineering Strong Python development experience Hands-on experience with LLM/GenAI evaluation Experience designing and implementing evaluation frameworks Experience developing custom or deterministic evaluators Experience integrating AI/ML quality checks into CI/CD Good understanding of LLM and AI agent architectures Nice to have Hands-on experience with AWS AgentCore Evaluation Experience with AWS Bedrock Guardrails, including PII detection Knowledge of CloudWatch metrics and production monitoring Experience with OpenTelemetry Familiarity with LangGraph, Strands Agents, or similar agent frameworks
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
What is Agentic Programming and Why Should Developers Care?
A 5-Step Open-Source Setup for Agentic Engineering
Introducing Redis Agent Memory Server
Never delegate the understanding