Senior AI Agent Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
We’re hiring a Senior AI Agent Engineer. You’ll build the agents in our product, and you’ll build the
evals that tell us whether each change made them better or worse.
The work is document-heavy rather than chat. The agents run multi-step, call tools, read long and
inconsistently formatted source material, check it against existing records, and produce output that a
person reviews before anything happens with it.
Accuracy matters more here than speed or novelty. Most of the engineering effort goes into precision,
traceability, and getting the agent to hand off to a human at the right moment.
You’ll report to the Head of Engineering and work with product and the full-stack team. If you’ve
shipped agents before, you’ve probably had the experience of changing a prompt and having no idea
whether you improved anything. That problem is most of this job.
What you’ll do
- Design and ship long-running, multi-step, tool-using agents on various AI SDKs and tooling,
included but not limited to the OpenAI Agents SDK, the Anthropic Agent SDK, the Vercel AI SDK,
LangGraph, MCP, and Temporal Cloud.
- Wrap our APIs and our partners’ APIs as tools an agent can call over MCP. Some of those systems
are old, single-tenant, and outside our control, so a fair amount of the work is translation.
- Get structured data out of long documents and match it against records that already exist. Expect
entity resolution and fuzzy matching, and expect much of it to run in batch.
- Stand up eval suites using various evaluation frameworks and tooling, included but not limited to
Promptfoo, Braintrust, LangSmith, DeepEval, LLM-as-judge methods, and custom harnesses.
Measure tool-use correctness, trajectory quality, and whether the agent finished the task.
- Every agent here produces a draft that a person signs off on. Build the citations and confidence
signals that make that review fast, and give the agent a clear way to escalate.
- Sit with product and domain experts and turn vague quality goals into something measurable.
Sometimes the only dataset available for that is tiny, or confidential, or both.
- Instrument production traffic, turn real customer interactions into golden datasets, and run them as
regression tests.
- Compare models against each other (OpenAI, Anthropic, open-weight), along with prompt
strategies and agent designs, and know what each option costs in latency and quality.
- Bootstrap quality signal for features that have no production traffic yet. That usually means
generating synthetic documents and test cases, including the ugly edge cases real customers will
Requirements
- 3+ years of engineering experience, including hands-on work on LLM or agent systems that real
users touched.
- You’ve evaluated agents, not only models, and you know why single-turn accuracy says little about
a multi-step run.
- You’ve integrated against APIs you don’t own, including old ones with bad documentation, and
turned them into something an agent can call reliably.
- You’re comfortable with document pipelines: pulling data out, normalizing it, and checking it against
a structured source of truth.
-
You’ve used at least one LLM evaluation framework, in-house tooling included.
-
You know how LLM-as-judge methods break down (position bias, verbosity bias, judge drift) and, * You’ve evaluated retrieval systems: RAG, hybrid search, reranking.
-
You’ve worked with agent orchestration frameworks like Temporal, LangGraph, or the OpenAI
Agents SDK, and you know how long-running tool use goes wrong.
-
You have a background in information retrieval or search relevance.
-
You’ve worked somewhere an agent’s output carried financial or compliance consequences.
-
You’ve built internal tooling that non-engineers used on their own to label and review model output. xkdbapo
About the company
Solicitante de Empleo
Iniciar sesión
Publica tu currÃculum, FirstIgnite makes software for university tech transfer offices. Those are the people who take research
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Loading talks and stories from around this role…