> Markdown version of [/jobs/ext/2970620-ai-engineer-harness](https://www.wearedevelopers.com/jobs/ext/2970620-ai-engineer-harness). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Engineer (Harness) - **Company:** Valarian Technologies Limited - **Location:** London, UK - **Salary:** £52,157.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Automated Storage and Retrieval Systems, Data Files, Memory Management, Statistical Hypothesis Testing, Python (Programming Language), Large Language Models, Multi-Agent Systems, Prompt Engineering, Kubernetes, Power Analysis (Cryptography), Data Analytics, Machine Learning Operations - **Published:** September 18, 2026 - **Apply:** https://www.adzuna.co.uk/jobs/details/5889031724 ## About the Role * Grounding in Statistics & Experimental Design: Demonstrated expertise in hypothesis testing, bootstrapping, power analysis, and statistical significance within probabilistic machine learning systems. * LLM-as-a-Judge Validation: Proven experience building, auditing, and validating LLM judge setups against human annotations, including tracking agreement metrics and mitigating judge biases. * Evaluation Dataset Engineering: Strong track record designing evaluation datasets, crafting precise scoring rubrics, managing label noise, and measuring inter-annotator agreement. * Evaluation of Non-Deterministic Multi-Step Systems: Deep comfort evaluating complex agent trajectories, tool execution sequences, latency, cost, and variance across repeated runs rather than relying solely on single-shot benchmarks. * Metric Design & Integrity: Exceptional judgment in defining metrics tied to real-world outcomes, with a keen eye for detecting benchmark gaming, target drift, or shortcut learning. * Rigor in Error Analysis: Ability to dissect complex execution logs, categorize failure modes, and synthesize clear, data-driven recommendations. * Hands-On LLM Integration: Practical proficiency with LLM APIs, prompt engineering, structured outputs, function/tool calling, and retrieval systems in Python. * Agent Architecture Awareness: Clear understanding of planning loops, tool orchestration, memory management, and multi-agent coordination, including their trade-offs and common failure modes. * Harness & Scaffolding Contribution: Ability to write clean, maintainable Python code to iterate on harness components such as guardrails, retry logic, state handling, and context management. Nice to have: * Experience with agent and evaluation frameworks (e.g., LangChain, LangGraph, AutoGen, lm-evaluation-harness, Promptfoo, or Ragas). * Exposure to running evaluation pipelines or agent workloads within secure, enclave, or Kubernetes environments. ## Description As an AI Harness Engineer at Valarian, you will own the experimental design, evaluation methodologies, and benchmark infrastructure for our AI models and autonomous agentic workloads. In an emerging domain with no standard playbook, you will serve as the bridge between rigorous data science, statistical validation, and production agent scaffolding. You will design robust evaluation datasets, architect calibrated LLM-as-a-judge pipelines, and build metrics that accurately capture multi-step agent performance under non-deterministic conditions. Your insights and error analyses will directly drive iterative improvements to our harness scaffolding, tool orchestration, and system safety guardrails. What you'll do: * Architect Evaluation Runtimes & Harnesses: Design and maintain scalable Python evaluation harnesses that measure task completion, trajectory quality, tool-use correctness, cost, latency, and variance across multi-step agentic systems. * Design Experiments & Validate Results: Apply rigorous statistical methods-including hypothesis testing, confidence intervals, bootstrapping, and power analysis-to determine whether performance changes are true system improvements or stochastic noise. * Build & Calibrate LLM-as-a-Judge Pipelines: Develop automated judging systems validated against human ground truth. Measure alignment using agreement metrics (Cohen's kappa, correlation, precision/recall) and systematically detect judge failure modes such as position bias, verbosity bias, self-preference, and prompt sensitivity. * Curate Benchmark Datasets & Rubrics: Define sampling strategies, detailed annotation guidelines, and scoring rubrics. Measure inter-annotator agreement and manage label noise to ensure benchmark integrity over time. * Deep Error & Trajectory Analysis: Conduct hands-on failure analysis on agent runs, cluster root causes into actionable error categories, and clearly communicate findings to both engineering teams and stakeholders. * Drive Harness Scaffolding Iteration: Translate evaluation results directly into architectural improvements across agent scaffolding, prompt design, tool execution loops, context window compaction, retries, and safety guardrails.