Senior Ai Agent Engineer

Firstignite
Toledo, Spain
18 days ago
Apply on www.buscojobs.com.es
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
3 years minimum
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Application Integration Architecture Automated Storage and Retrieval Systems Cloud Computing Data Normalization Information Retrieval Regression Testing Next.js Large Language Models Multi-Agent Systems Virtual Agents

Job description

About FirstIgniteFirstIgnite makes software for university tech transfer offices. Those are the people who take researchcoming out of a university lab and get it patented, licensed, or spun out into a company.The roleWere hiring a Senior AI Agent Engineer. Youll build the agents in our product, and youll build theevals that tell us whether each change made them better or worse.The work is document-heavy rather than chat. The agents run multi-step, call tools, read long andinconsistently formatted source material, check it against existing records, and produce output that aperson reviews before anything happens with it.Accuracy matters more here than speed or novelty. Most of the engineering effort goes into precision,traceability, and getting the agent to hand off to a human at the right moment.Youll report to the Head of Engineering and work with product and the full-stack team. If youveshipped agents before, youve probably had the experience of changing a prompt and having no ideawhether you improved anything. That problem is most of this job.What youll doDesign and ship long-running, multi-step, tool-using agents on various AI SDKs and tooling, included but not limited to the OpenAI Agents SDK, the Anthropic Agent SDK, the Vercel AI SDK, LangGraph, MCP, and Temporal Cloud.Wrap our APIs and our partners APIs as tools an agent can call over MCP. Some of those systems are old, single-tenant, and outside our control, so a fair amount of the work is translation.Get structured data out of long documents and match it against records that already exist. Expect entity resolution and fuzzy matching, and expect much of it to run in batch.Stand up eval suites using various evaluation frameworks and tooling, included but not limited to Promptfoo, Braintrust, LangSmith, DeepEval, LLM-as-judge methods, and custom harnesses. Measure tool-use correctness, trajectory quality, and whether the agent finished the task.Every agent here produces a draft that a person signs off on. Build the citations and confidence signals that make that review fast, and give the agent a clear way to escalate.Sit with product and domain experts and turn vague quality goals into something measurable. Sometimes the only dataset available for that is tiny, or confidential, or both.Instrument production traffic, turn real customer interactions into golden datasets, and run them as regression tests.Compare models against each other (OpenAI, Anthropic, open-weight), along with prompt strategies and agent designs, and know what each option costs in latency and quality.Bootstrap quality signal for features that have no production traffic yet. That usually means generating synthetic documents and test cases, including the ugly edge cases real customers will eventually send us, and knowing where synthetic data stops being a good proxy.Write the templates, docs, and tooling the rest of the team needs to run evals without coming to you.Required Qualifications:3+ years of engineering experience, including hands-on work on LLM or agent systems that real users touched.Youve evaluated agents, not only models, and you know why single-turn accuracy says little about a multi-step run.Youve integrated against APIs you dont own, including old ones with bad documentation, and turned them into something an agent can call reliably.Youre comfortable with document pipelines: pulling data out, normalizing it, and checking it against a structured source of truth.Youve used at least one LLM evaluation framework, in-house tooling included.You know how LLM-as-judge methods break down (position bias, verbosity bias, judge drift) andwhat to do about it.You can tell a real regression from noise, and design an experiment that answers the question being asked instead of a nearby one.You can read a customer call transcript, work out which failures matter, and ship a fix and an eval for them.You write clearly. Engineers wont act on eval results they dont read or dont trust.Youre based somewhere between New York time (ET) and Western European time. Italy is thefurthest east we can go.You might currently be titledAI Agent Engineer · Agentic AI Engineer · Senior AI Engineer · AI Systems Engineer · Lead AIDeveloper · AI Solutions Architect · AI Integration Specialist · Senior Data Scientist, AITitles are all over the place in this space. If the work above matches what you already do, apply. Wellgo on what youve shipped.Preferred Qualifications:Youve evaluated retrieval systems: RAG, hybrid search, reranking.Youve worked with agent orchestration frameworks like Temporal, LangGraph, or the OpenAIAgents SDK, and you know how long-running tool use goes wrong.You have a background in information retrieval or search relevance.Youve worked somewhere an agents output carried financial or compliance consequences.Youve built internal tooling that non-engineers used on their own to label and review model output.This is a fully remote, full-time permanent position available to candidates located within the New York (ET) through Western Europe time zones, with flexible working hours to support collaboration across regions.

Requirements

3+ years of engineering experience, including hands-on work on LLM or agent systems that real users touched. Youve evaluated agents, not only models, and you know why single-turn accuracy says little about a multi-step run. Youve integrated against APIs you dont own, including old ones with bad documentation, and turned them into something an agent can call reliably. Youre comfortable with document pipelines: pulling data out, normalizing it, and checking it against a structured source of truth. Youve used at least one LLM evaluation framework, in-house tooling included. You know how LLM-as-judge methods break down (position bias, verbosity bias, judge drift) and, Youve evaluated retrieval systems: RAG, hybrid search, reranking. Youve worked with agent orchestration frameworks like Temporal, LangGraph, or the OpenAI Agents SDK, and you know how long-running tool use goes wrong. You have a background in information retrieval or search relevance. Youve worked somewhere an agents output carried financial or compliance consequences. Youve built internal tooling that non-engineers used on their own to label and review model output. This is a fully remote, full-time permanent position available to candidates located within the New York (ET) through Western Europe time zones, with flexible working hours to support collaboration across regions.

About the company

About FirstIgnite FirstIgnite makes software for university tech transfer offices. Those are the people who take research coming out of a university lab and get it patented, licensed, or spun out into a company.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.buscojobs.com.es
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:31 min

Introduction to the GraphQL, Apollo, and Next.js stack

Josh Goldberg · JS Congress

3:21 min

Navigating the ethical integration of virtual colleagues

Rudi Bauer Rudi Bauer +1 · Cappuccino with HR

2:27 min

Introduction to WebAssembly in a cloud computing context

Edo Edo · World Congress 2024

4:37 min

Transitioning to a hybrid human and agent workforce

Thomas Dohmke Thomas Dohmke · World Congress 2025

40 sec

Empowering automated workflows with agentic AI models

Mike Mike · World Congress 2025

3:06 min

Leveraging the Next.js framework for performance improvements

Eileen Fürstenau Eileen Fürstenau · World Congress 2024

Videos

See all

Related articles

See all