> Markdown version of [/jobs/ext/2018608-ai-engineer](https://www.wearedevelopers.com/jobs/ext/2018608-ai-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Engineer - **Company:** Pennylane Sas - **Location:** Spain (Remote available) - **Contract:** Permanent contract - **Skills:** A/B Testing, Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Audit Trail, Common ISDN Application Programming Interface (CAPI), Financial Software, Python (Programming Language), Machine Learning, Regression Analysis, Role-Based Access Control, Ruby, Software Engineering, Datadog, Large Language Models, AI Platforms, Kubernetes, Optimization Algorithms, Deployment Automation, Machine Learning Operations - **Published:** August 11, 2026 - **Apply:** https://es.trabajo.org/oferta-4112-6d7fbffc2fad919664241ea320b5ed40 ## About the Role Are you looking to have an impact on the daily life of millions of entrepreneurs in France (and tomorrow in Europe)? Are you looking for a work environment that values trust, proactivity, and autonomy? Are our Engineering principles aligned with your vision? Then Pennylane is the right place for you Our vision We aim to become the most beloved financial Operating System of French SMEs and Accounting Firms (and soon, European ones). We help entrepreneurs rid themselves of time-consuming tasks related to accounting and finance while giving them access to the key financial information they need to make the best decisions for their business. About Us Pennylane is one of the fastest growing Fintechs in France (and soon in Europe). In 5 years of existence, we've managed to: Make ourselves known as a groundbreaking accounting and financial software for small businesses and their accountants Raise a total of €400 million, including from Sequoia - the famous Silicon Valley fund that, architectures, multimodal AI, and optimization techniques). WHAT You Can Expect From Your Life At Pennylane Within one month: You will learn everything about our company, our teams, and our vision during the first onboarding week. You will get familiar with our stack and AI tooling (agent harness, tool registry & MCP, evaluation and observability tooling, model gateway), and have delivered a few small projects which will give you a concrete taste of our tools & processes. You will be given time to meet your future stakeholders, and gain a deep knowledge of our product and operations. Within 3 months: You will be fully in charge of items in our roadmap, defining and prioritizing your tasks autonomously, and will own an agentic use case end to end in production, together with its evaluation set and quality metrics. You will be comfortable with our technical stack (Python, agentic frameworks & MCP, evaluation and observability tooling, Kubernetes & AWS). You will, first-class discipline: golden datasets, LLM-as-judge, human eval, A/B testing, and measuring agent quality, regressions and edge cases. Have a good grasp of applied LLM / ML and AI infrastructure (model serving, vector databases, cost and latency). Nice to have: model fine-tuning and post-training (SFT, DPO, RL), MCP, and familiarity with Ruby. Have a balanced blend of technical, business and product skills, communicate well (including with non-technical domain experts), and are fluent in English (French is not mandatory). What does the recruitment process look like? A first interview with our Talent Acquisition Manager A case study interview to discuss a topic closely related to one of our priorities (75 min) A past-project interview to hear about your experience (60 min) An interview with our Tech & Product leaders to discuss our company culture (60 min) What we do to make your work life easier Wherever you are based, you'll get 25 vacation days paid ## Description agent interface that makes Pennylane readable and actionable by AIs (tools, MCP), and the Pennylane orchestrator agent itself (Studio Assistant). Models are commoditizing fast. What makes the difference is everything around them: the harness, the context, the tools, the guardrails. And the evaluation that proves it works on real accounting workflows. That is exactly what this role owns. By joining us as an AI Engineer, you will have a pivotal role in large projects, bringing your software engineering and applied LLM / agentic systems expertise to help us meet our high delivery standards, at scale and with the level of trust our users require. HOW you will contribute to the company as an AI Engineer As an AI Engineer, you will be part of one of our ML & AI teams (25+ people, growing from 2 to 5 teams by the end of 2026), depending on your profile and our needs: AI Capabilities & Infrastructure (CAPI) : the shared foundations every squad builds on: agent harness, agentic runtime & sandboxing, model gateway, tool registry & MCP, guardrails, observability, LLM FinOps, and reusable AI capabilities. Agent Behaviour : making our agents behave correctly, safely and measurably: behaviour steering, domain tuning, evaluation methodology & harness, and later model fine-tuning. Studio Assistant : the financial orchestrator agent for accountants and SMEs: understanding intent, picking tools, launching jobs, producing reports, asking for validation when needed. Whichever team you join, the core of the job is the same: You will build and harden agentic loops in production : LLM call orchestration, tool calling, streaming & events, error recovery, checkpointing and resumability. With security, permissions and cost in mind (agentic RBAC, audit trail, isolated execution, token and latency budgets). You will own context & memory engineering: context construction, retrieval, compaction, short- and long-term memory, and the latency / cost trade-offs that come with them. You will make Pennylane usable by agents, turning business capabilities into well-specified, documented, versioned and tested tools exposed internally and through MCP, and steering agent behaviour (system prompts and instructions, skills, planning and tool-selection strategies, guardrails, anti-prompt-injection). You will treat evaluation as a first-class discipline: golden datasets, LLM-as-judge, human eval, regression tracking and error analysis. Turning production failures into systematic improvements. You will collaborate with Product teams and domain experts (accountants) on live use cases : ComptAssistant, Studio Assistant, MCP, document extraction, Autopilot bookkeeping, revision. To make sure that fulfilling real user needs is always at the center of what we do. Stay Ahead of the Curve: Continuously monitor the AI landscape for emerging trends, State-of-the-Art (SOTA) models, and breakthrough research (e.g., new LLM, contribute to larger cross-team projects. Within 6 months: You will proactively contribute to the team's roadmap. You will work with engineers and data practitioners on improving our agent harness, our evaluation practices and our AI platform. You will share your learnings and best practices within the team. And beyond: the ML & AI teams will continue growing with the company Which means: Opportunities to recruit and mentor new team members, Increased accountability in project leadership, Responsibilities to design and implement new processes, tools and best practices to make sure that your team works even more efficiently. Who are we looking for? You're the right candidate if you: Have 5-8 years of experience and are very strong in Python Hands-on experience building LLM and agentic systems at scale in production: prompting, tool use, context construction, RAG, and handling failure, state and reliability (not just calling a model API)., Treat evaluation as a ## Related Videos - [Agentic employees in world's most downloaded FinTech app](https://www.wearedevelopers.com/videos/100123-agentic-employees-in-world-s-most-downloaded-fintech-app) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [Coffee with Developers: David Heinemeier Hansson](https://www.wearedevelopers.com/videos/875-coffee-with-developers-david-heinemeier-hansson) - [Inside Bitpanda's Tech Stack: Scaling a European Fintech Leader - Markus Dorner](https://www.wearedevelopers.com/videos/1979-inside-bitpanda-s-tech-stack-scaling-a-european-fintech-leader-markus-dorner) - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/videos/100333-instant-kai-sandboxes-with-vcluster-multi-tenant-multi-scheduler-gpu-sharing) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Best Companies to Work For in Paris: Top 25 Companies in 2023 ](https://www.wearedevelopers.com/magazine/190-best-companies-to-work-for-in-paris-top-25-companies-in-2023) - [Dev Digest 132 - Binging WADFlix?](https://www.wearedevelopers.com/magazine/473-dev-digest-132-binging-wadflix)