Applied AI Engineer, AI Platform
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+8 more
Job description
- Evaluate and improve agent performance. Build the evaluation layer: test cases built from real documents with the output we expect, regression suites in CI, human review where correctness is non-negotiable, and clear success criteria for an agent completing a complex task end to end. Then move the numbers that matter - accuracy, latency, cost.
- Own the AI application architecture. Orchestration and multi-agent design, tool contracts, memory, and context engineering (RAG, MCP) with clear domain boundaries - plus the guardrails, approvals, and human-in-the-loop controls that anything touching investor money requires.
- Run it in production. Versioned, feature-flagged rollout of agent versions; rate limits, provider fallback, and regional failover within EU data residency; and the observability to trace a failure across services and turn it into a fix rather than a theory.
- Make it a platform, not a project. Document parsing and extraction consolidates into this team, and shared evaluation and observability become something other teams consume rather than rebuild. You set the patterns other engineers inherit, partner closely with DevX, and mentor engineers across teams on agent and LLM practice.
Your First 90 Days
- Review our MCP offering end to end and ship it to the first customers, with the access boundaries, evaluation and observability a customer-facing surface needs.
- Stand up shared observability for our AI workloads - token usage, estimated cost, latency, failures and retries per provider - on a dashboard people actually open during an incident.
- Publish v1 of our agent patterns (domain boundaries, tool contracts, prompt and eval conventions), reviewed with DevX and adopted by at least one team outside AI Platform., * Frontend: TypeScript, Svelte and React
- Backend: Node.js, Nest.js
- Database: MySQL, PostgreSQL
- AI: Mastra, Vercel ai-sdk, Google Vertex (Gemini) with AWS Bedrock fallback in EU regions
- Infrastructure: Kubernetes on AWS
- Observability: Datadog
- Auth & Internal tools: FusionAuth, Retool, 2. Hiring Manager Interview (60 min) - Explore product-focused culture, collaboration, ownership, self-awareness, and quality standards. 3. Technical Interviews (2 x 60 min) - Two technical conversations, one with our engineers to discuss topics like system and agent-design questions, and one focused on core AI technical topics 4. Final Round with our CTO Leandro (45 min) - Discuss ways of working, bunch’s engineering vision, and team culture.
Questions? Reach out to Recruiting@bunch.capital
About bunch
bunch is building the operating infrastructure for the next generation of private markets. We combine AI-powered automation with deep regulatory expertise to replace fragmented spreadsheets and manual processes with one integrated platform across the fund lifecycle, purpose-built for private markets heading toward $32 trillion in Assets Under Management.
We’ve 4x our ARR in 2025, crossed 150 fund managers and 12,000 LPs on the platform, and just closed our $35M Series B in May 2026. We’re looking for ambitious people who want real ownership of hard problems, and who care about building infrastructure that actually matters to the people using it.
Requirements
- Experience: 5+ years building production software, including at least one agent or LLM-powered capability you took end to end and still owned once it was live
- Agents: real depth in orchestration and context engineering - tool contracts, memory, RAG, multi-agent design - with Mastra, ai-sdk, LangGraph or similar. You’ve integrated agents into a real product and its authorization model, not built prototypes for someone else to productionise
- Evaluation: you know how to make a non-deterministic system measurable, from test cases built on real data to regression suites and human review, and you can tell the difference between a model that improved and a benchmark that got easier
- Production: you own what you ship. Reading traces, diagnosing failure modes, tuning cost and latency, handling provider rate limits and fallback without drama
- Backend: You have experience with TypeScript and/or Python.
- Platform mindset: you build for other engineers as much as for end users, and the standards you set get adopted because people trust you, not because they are written down
- Pragmatic: you start from the business outcome, choose the deterministic solution when it’s the right one, and know when good enough is good enough
- Experience in fintech, private markets, or another regulated, document-heavy domain is a plus, as is working under EU data-residency constraints
Benefits & conditions
- Customizable benefits package (wellbeing, sport, mobility, food, and more)
- 28 days of vacation, plus 2 company days and local public holidays
- Hybrid setup (3 days/week in office)
- Up to 6 remote calendar weeks a year
- A great tech and work setup
- Work with a diverse team of 130+ bunchies from 40+ countries, with leaders who are best-in-class in their domains
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path
How to Become an AI Engineer
Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?
MLOps And AI Driven Development