> Markdown version of [/jobs/ext/1997062-founding-ai-engineer](https://www.wearedevelopers.com/jobs/ext/1997062-founding-ai-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Founding AI Engineer - **Company:** Mondrio - **Location:** Amsterdam, Netherlands - **Experience:** Expert - **Salary:** €10,833.0 - €13,750.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Business Logic, Code Review, Cursor (Graphical User Interface Elements), Github, Python (Programming Language), MongoDB, Systems Development Life Cycle, Next.js, Software Factory, TypeScript, ReactJS, Large Language Models, Fastapi - **Published:** August 9, 2026 - **Apply:** https://www.adzuna.nl/details/5833604242 ## About the Role * 8+ years of engineering experience with strong and recent production LLM depth. * You have shipped LLM-powered product features to production and owned them after launch. * You have built evals and observability for LLM systems yourself. Running someone else's dashboard does not count. * Strong communication skills to bridge the technical gap around non-deterministic engineering to less savvy clients and partners * Product engineer instincts: you pick your own scope and choose the pragmatic option over the interesting one. This is not a research-lab role. * You can show how AI coding tools fit into your work today. We weigh that over where you studied or previous role. Nice to have: * Experience with MCP or building tools for LLM agents. * Experience with platforms such as LangChain, LlamaIndex, Braintrust, OpenRouter * Work in a domain where correctness is audited, such as pricing, billing, or payments. * Familiarity with data residency or compliance constraints. SOC2 and GDPR shape what you build against. Client: Typescript, React, Vercel Chat SDK Server: Python, FastAPI Data: Mongo, Atlas ## Description Mondrio's AI recommends prices, and expert Pricing Architects stay in the loop on the high-stakes calls. Your mandate is to build the evaluation systems and feedback loops that let the AI earn more of that trust. This is applied AI on a problem where quality is measurable in customer revenue. What you'll do * Build evaluation for AI pricing recommendations: eval harnesses and benchmarks that use tracked pricing outcomes as ground truth. Expert review is manual today, and you make it systematic. * Take AI personas further. They simulate B2B buying committees and behavioral effects such as new versus existing customers, grounded in usage data and call transcripts. Automate the parts of persona training that are still manual. * Own LLM infrastructure: routing across current (Anthropic and Google) and future models, with explicit cost, latency, and quality tradeoffs. * Maintain infra and data residency boundaries (e.g. model calls for EU customers must remain within the EU) as we add providers and scale up operations. * Extend the MCP server that LLM agents, including our customers' own agents, use to drive the platform. A feature is done when an agent can drive it through MCP, not when the React component renders. * Work within our typed ontology of pricing entities (Pydantic models for SKU, Proposition, Persona, Quote) so model outputs land in structured, auditable form. Your first 90 days First 30 Days: Foundation & Guardrails * Model routing across Anthropic and Google has explicit cost and latency budgets, and the fail-closed EU residency guarantee covers every model call. * Pinpoint systemic latency, data drift, or cold-start issues in the continuous pricing loop. * Baseline current prompt and model outputs against our typed ontology to prepare for release-gating evals. * Conduct code reviews and lead a technical session on advanced AI/ML patterns for the team. By Day 60: Trust & Automation * An eval harness runs on every model or prompt change, and the team trusts its benchmarks enough to gate a release on them. * Manual steps in persona training now run as an automated pipeline built on the same usage data and call transcripts. By Day 90: Closed-Loop Impact * Tracked pricing outcomes feed back into recommendation quality, so evals measure revenue impact, not proxy scores. * Participate actively in interview loops to scale the engineering team and mentor mid-level engineers., Security: SOC2 Type 1/Type 2, GDPR compliant, EU and US data residency We are under ten people, and everyone ships and talks to customers. Engineers are product engineers: you own outcomes, scope your own work, and demo every week. We appreciate fast feedback loops, real collaboration, and a team that genuinely enjoys working together. If that's how you do your best work, you'll fit right in. There is no separate PM layer. Our SDLC is AI-native as we move to a robust software factory. Claude Code and Cursor are standard kit, and LLM agents write and review pull requests behind automated review gates. We keep pull requests small, around 500 lines, because the research on small batches holds up and we act on it. The architecture rules are short and enforced: the API is the product so it evolves additively, business logic stays server-side, and writes are audited by default for compliance. Your evals are how we stay honest about whether the AI actually works., Send a short note about something relevant you've built, plus frustrations around B2B pricing and your best take on how it should work instead. If you have public work (GitHub, writing, a talk), link it. If your best work is private, which most is, a paragraph walking us through it counts just as much. ## Related Videos - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [40 Minutes to Build a Serverless COVID-19 REST and GraphQL APIs](https://www.wearedevelopers.com/videos/208-40-minutes-to-build-a-serverless-covid-19-rest-and-graphql-apis) - [GraphQL + Apollo + Next.js: A Lovely Trio](https://www.wearedevelopers.com/videos/311-graphql-apollo-next-js-a-lovely-trio) - [Coffee with Developers - Maria Apazoglou](https://www.wearedevelopers.com/videos/1209-coffee-with-developers-maria-apazoglou) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) - [Building AI Applications with LangChain and Node.js](https://www.wearedevelopers.com/videos/1512-building-ai-applications-with-langchain-and-node-js) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [Dev Digest 210: AI Agents Are Go! Is MCP Dead? LLMs Crack Anonymity](https://www.wearedevelopers.com/magazine/709-dev-digest-210-ai-agents-are-go-is-mcp-dead-llms-crack-anonymity) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 132 - Binging WADFlix?](https://www.wearedevelopers.com/magazine/473-dev-digest-132-binging-wadflix)