> Markdown version of [/jobs/ext/1488238-ai-engineer](https://www.wearedevelopers.com/jobs/ext/1488238-ai-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Engineer - **Company:** Discovered - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, Continuous Integration, Software Debugging, Python (Programming Language), OAuth, Recommender Systems, Software Engineering, TypeScript, ReactJS, Large Language Models, Kubernetes, Data Lineage, Terraform, Pagination, Data Pipelines, Human in the Loop - **Published:** July 29, 2026 - **Apply:** https://app.dover.com/apply/Discovered/d3c07480-cf92-4c42-9468-164c0439e175 ## About the Role A first-principles thinker. You understand why things work, not just how. You can go five levels deep on eval design, prompt architecture, and where to put the human in the loop. Always improving. You're not satisfied with "good enough." You actively seek ways to get better at your craft and make systems better over time. Requirements 5+ years in software engineering, with meaningful recent time on LLM-backed or ML-backed production systems, including 2+ owning a production system end to end. Python, React, Typescript and strong systems fundamentals. You write production services, not notebooks. LLM application engineering in production. Agents, tool use, structured outputs, retrieval, prompt architecture. You've shipped something real that used them and stayed up. Evaluation systems. You've built eval suites for non-deterministic systems: golden datasets, regression gates, offline and online scoring. You've thought hard about what "correct" means when there are many correct answers. Observability for AI systems. Tracing, run inspection, cost and token accounting. Langfuse, LangSmith, Braintrust, OpenTelemetry or equivalent., * Browser automation or authenticated web agents (Playwright, Puppeteer, computer-use models) * Applied ML: ranking, scoring, classification, or recommendation systems in production * Python, React, Typescript * Terraform, Kubernetes * Prior experience at a fast-moving startup ## Description That means evals that catch regressions before they ship, observability that explains why an agent did what it did, provenance that traces every claim back to its source, and interfaces that let a non-engineer SEO strategist run and review the work. You'll build the harnesses and eval suites that gate agent changes in CI, the lineage and step-trace tooling that makes agent reasoning inspectable, and the surfaces the team actually uses. Then you'll close the loop: instrument the output, feed the signal back, and make the agents measurably better every cycle. You'll also extend what agents can reach. Today they work over APIs and our own data. Next they operate real websites along withsession handling, and navigating interfaces that were never designed for machines. The hard problem is reliability and legibility. An agent that's right 80% of the time and can't tell you which 80% is worthless. An agent that's right 95% of the time, shows its working, and flags its own uncertainty is a product. You report to the CTO and work alongside the Automations lead, with our AI & Data team owning the platform beneath you. You own your evals, your CI, and your monitoring. What You'll Do Agent eval harnesses. Golden suites, deterministic checks, and online scoring that gate every agent change in CI. A regression should fail a build, not a client report. Agent observability. Step traces, token and cost accounting, run-level SLOs, and failure taxonomies. You'll know the difference between "the agent ran" and "the agent was right." Data lineage and provenance. Every claim an agent makes should be traceable to its source. Which inputs, which tool calls, which reasoning steps produced this output. End-to-end agents with human review. Agents that complete real work start to finish, with a curated review surface so an expert can approve, correct, or reject, and so that correction becomes training signal. Interfaces for non-engineers. Dashboards and controls that let the SEO team run, inspect and trust agent work without asking an engineer. Expanding what agents can do. Agents that open pull requests against real repositories, authenticate through OAuth, and operate real sites and CMSs through the browser. Every new capability is a new class of work the swarm can take on. Closing the loop. Turn signals we already collect, such as AI perception and citation data, into concrete client value: strategy recommendations, prioritised actions, and measurable outcomes. Algorithms and scoring models. Not everything should be a prompt. You'll build and tune the scoring systems behind what we recommend: internal linking and semantic relevance scoring, Reddit opportunity and thread scoring, content and citation quality. These are ranking and classification problems with real feedback data behind them, and where a model beats a prompt, you build the model. Shared, composable modules. Reusable capability blocks that both agents and workflows compose, so an improvement lands everywhere at once instead of being reimplemented per agent. Scaling across clients. Take a workflow that works for one client and make it run reliably for many, with per-client configuration, isolation, and failure that stays contained. The Ideal Person for This Role A builder who ships. You care about getting working systems into production, not endless planning or polish. You've built AI systems people actually rely on. An operator, not just an architect. You don't just design systems, you run them. You find satisfaction in making things reliable, not just making them work once in a demo. An owner. You take responsibility for outcomes, not just tasks. When an agent silently produces bad output, you catch it, fix it, and build the eval that stops it recurring. Maniacal about detail. This is the one that matters most here. AI has made it trivially cheap to produce a large volume of plausible-looking work, and most of that work is slop. We are building the opposite: output that holds up when a client reads it line by line. If you'd rather ship one thing that is provably right than ten that are probably fine, this is your team. If you cannon volume and let the reviewer sort it out, it isn't. Sceptical of your own output. You assume the model is wrong until measured. You've been burned by a demo that worked and a production run that didn't, and you build accordingly. Humble and curious. You acknowledge what you don't know, ask good questions, and genuinely want to learn. You take feedback as a gift, not a threat., Third-party API integration. Auth flows, rate limits, pagination, breaking changes. Not just calling endpoints, but handling the full operational reality. Own your infrastructure. Containers, CI/CD, deployment, monitoring, credential management. No platform team to hand off to. Product sense for expert users. You've built tooling that domain experts, not engineers, use daily. You know that an unexplained AI output is an unusable one. Collaborative. You'll define contracts with the engineers who own the infra beneath you. You document decisions, write clear specs, and communicate tradeoffs in writing. ## Related Videos - [Watch Tests Go Brrrr! : Getting Started with Cypress in ReactJS](https://www.wearedevelopers.com/videos/282-watch-tests-go-brrrr-getting-started-with-cypress-in-reactjs) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Keeping applications secure by evolving OAuth 2.0 and OpenID Connect](https://www.wearedevelopers.com/videos/100152-keeping-applications-secure-by-evolving-oauth-2-0-and-openid-connect) - [You are not an AI developer](https://www.wearedevelopers.com/videos/1148-you-are-not-an-ai-developer) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) - [LLMs in the wild: Building an AI agent that survives production](https://www.wearedevelopers.com/videos/100319-llms-in-the-wild-building-an-ai-agent-that-survives-production) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Never delegate the understanding](https://www.wearedevelopers.com/magazine/749-never-delegate-the-understanding) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)