Founding Engineer, Agent Systems

TechTree
Greater London, UK
2 days ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Starter
Experience required
1 year minimum
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Regression Testing TypeScript Management of Software Versions Large Language Models Multi-Agent Systems

Job description

A seed-stage company is building agent-native risk infrastructure: risk management and trust building, delivered by AI agents, for a world increasingly run by them. Their agents sit on top of proprietary data and reassess continuously rather than at fixed checkpoints - so customers spend their time deciding and acting on what matters, not assembling evidence to get there.

Seven-figure revenue within months of launch, on multi-year contracts with leading enterprises in financial services, regulated technology, and healthcare. Founders from Palantir, Oxford, Stanford, and ETH. Backed by leading UK and US institutional investors and angels from Meta, Isomorphic Labs, Palantir, and SpaceX.

The role

You own the agent platform: the orchestration, evals, and reliability work that turns model calls into product features customers trust. The bar isn’t that the demo works - it’s that a domain expert reading the agent’s output considers it at the level of a peer.

This isn’t a research role at its core: the team consumes frontier APIs and makes them production-grade. They push them hard - hard enough to have recently found and reported a bug in the Anthropic API that took their engineers weeks to reproduce. At that level, the line between using models and studying them gets thin, so if research-flavoured work pulls at you, there’s room to follow it.

What you’ll do

Evals for fuzzy, high-stakes outputs: assessments, policy interpretation, control mapping

Reliability infrastructure: retries, fallbacks, circuit breakers, prompt versioning

Set the internal standard for what “good enough to ship” means for AI features

What you bring

Backend engineering in TypeScript (or comparable), with 1-2+ years shipping production LLM features

Experience with agent frameworks, tool calling, and multi-step orchestration

Production evals: dataset curation, LLM-as-judge failure modes, regression testing under model swaps

Strong systems thinking: async, queues, idempotency

Comfort being the named owner of AI quality, including saying no when needed

Nice to have: Anthropic, OpenAI, or open-weight APIs in production at scale prompt-injection or agent-security work background in compliance, audit, or any domain where correctness is fuzzy and stakes are high

Working here

King’s Cross, London (Gridiron building) - in-person by default, flexibility for days that need it

Daily team lunch, specialty coffee, roof terrace, on-site showers, serious AI tooling and API budgets

Three-stage interview: behavioural phone screen, technical phone screen, paid on-site work trial - under two weeks from first conversation

Requirements

Backend engineering in TypeScript (or comparable), with 1-2+ years shipping production LLM features

Experience with agent frameworks, tool calling, and multi-step orchestration

Production evals: dataset curation, LLM-as-judge failure modes, regression testing under model swaps

Strong systems thinking: async, queues, idempotency

Comfort being the named owner of AI quality, including saying no when needed

Nice to have: Anthropic, OpenAI, or open-weight APIs in production at scale prompt-injection or agent-security work background in compliance, audit, or any domain where correctness is fuzzy and stakes are high

About the company

A seed-stage company is building agent-native risk infrastructure: risk management and trust building, delivered by AI agents, for a world increasingly run by them. Their agents sit on top of proprietary data and reassess continuously rather than at fixed checkpoints - so customers spend their time deciding and acting on what matters, not assembling evidence to get there.

Seven-figure revenue within months of launch, on multi-year contracts with leading enterprises in financial services, regulated technology, and healthcare. Founders from Palantir, Oxford, Stanford, and ETH. Backed by leading UK and US institutional investors and angels from Meta, Isomorphic Labs, Palantir, and SpaceX.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:00 min

Misconceptions about TypeScript safety capabilities

Simone Sanfratello ¡ JS Congress

3:32 min

Building evaluation frameworks for automated regression testing

Andreas Erben Andreas Erben ¡ World Congress 2025

2:26 min

Comparing single-shot prompts and multi-agent systems

Dr. Alexander Wachtel Dr. Alexander Wachtel +1 ¡ World Congress 2025

2:51 min

Applying multi-agent AI debate to everyday engineering workflows

Lior Schejter Lior Schejter ¡ Europe 2026 Virtual

2:55 min

Gaining TypeScript benefits using JSDoc alternatives

Simone Sanfratello ¡ JS Congress

1:24 min

Building client-facing AI agents for engineering teams

Alfonso Graziano Alfonso Graziano ¡ Coffee With Developers

Videos

See all

Related articles

See all