Founding Engineer, Agent Systems
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
A seed-stage company is building agent-native risk infrastructure: risk management and trust building, delivered by AI agents, for a world increasingly run by them. Their agents sit on top of proprietary data and reassess continuously rather than at fixed checkpoints - so customers spend their time deciding and acting on what matters, not assembling evidence to get there.
Seven-figure revenue within months of launch, on multi-year contracts with leading enterprises in financial services, regulated technology, and healthcare. Founders from Palantir, Oxford, Stanford, and ETH. Backed by leading UK and US institutional investors and angels from Meta, Isomorphic Labs, Palantir, and SpaceX.
The role
You own the agent platform: the orchestration, evals, and reliability work that turns model calls into product features customers trust. The bar isnât that the demo works - itâs that a domain expert reading the agentâs output considers it at the level of a peer.
This isnât a research role at its core: the team consumes frontier APIs and makes them production-grade. They push them hard - hard enough to have recently found and reported a bug in the Anthropic API that took their engineers weeks to reproduce. At that level, the line between using models and studying them gets thin, so if research-flavoured work pulls at you, thereâs room to follow it.
What youâll do
Evals for fuzzy, high-stakes outputs: assessments, policy interpretation, control mapping
Reliability infrastructure: retries, fallbacks, circuit breakers, prompt versioning
Set the internal standard for what âgood enough to shipâ means for AI features
What you bring
Backend engineering in TypeScript (or comparable), with 1-2+ years shipping production LLM features
Experience with agent frameworks, tool calling, and multi-step orchestration
Production evals: dataset curation, LLM-as-judge failure modes, regression testing under model swaps
Strong systems thinking: async, queues, idempotency
Comfort being the named owner of AI quality, including saying no when needed
Nice to have: Anthropic, OpenAI, or open-weight APIs in production at scale prompt-injection or agent-security work background in compliance, audit, or any domain where correctness is fuzzy and stakes are high
Working here
Kingâs Cross, London (Gridiron building) - in-person by default, flexibility for days that need it
Daily team lunch, specialty coffee, roof terrace, on-site showers, serious AI tooling and API budgets
Three-stage interview: behavioural phone screen, technical phone screen, paid on-site work trial - under two weeks from first conversation
Requirements
Backend engineering in TypeScript (or comparable), with 1-2+ years shipping production LLM features
Experience with agent frameworks, tool calling, and multi-step orchestration
Production evals: dataset curation, LLM-as-judge failure modes, regression testing under model swaps
Strong systems thinking: async, queues, idempotency
Comfort being the named owner of AI quality, including saying no when needed
Nice to have: Anthropic, OpenAI, or open-weight APIs in production at scale prompt-injection or agent-security work background in compliance, audit, or any domain where correctness is fuzzy and stakes are high
About the company
A seed-stage company is building agent-native risk infrastructure: risk management and trust building, delivered by AI agents, for a world increasingly run by them. Their agents sit on top of proprietary data and reassess continuously rather than at fixed checkpoints - so customers spend their time deciding and acting on what matters, not assembling evidence to get there.
Seven-figure revenue within months of launch, on multi-year contracts with leading enterprises in financial services, regulated technology, and healthcare. Founders from Palantir, Oxford, Stanford, and ETH. Backed by leading UK and US institutional investors and angels from Meta, Isomorphic Labs, Palantir, and SpaceX.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role â technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
What is Agentic Programming and Why Should Developers Care?
Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?
A 5-Step Open-Source Setup for Agentic Engineering
Dev Digest 137 - AI'm not sure about this