AI Engineer for BASF's DevHub
Role details
Job location
Tech stack
Job description
WELCOME TO BASF No deje pasar esta oportunidad, inscríbase rápidamente si su experiencia y habilidades coinciden con lo que se indica en la siguiente descripción. Digitalization is a true part of BASF's DNA - creating new customer experiences, driving business growth, and making processes more efficient. Global Digital Services drives BASF's digital transformation through innovative, global, high-quality digital products and a strong agile culture, and the Digital Hub Madrid is one of our key global delivery locations. We are seeking a hands-on AI Engineer for BASF's DevHub - the Internal Developer Platform (IDP) used by thousands of engineers and product teams across BASF. DevHub already ships an enterprise AI Gateway (50+ governed models, Entra ID, EU data residency, per-cost-center billing, Grafana observability) and a catalog that is a schema-validated knowledge graph of every product and its infrastructure. Your mission is to make AI a first-class platform capability: build reusable, production-grade AI services and developer experiences that help users discover, create, configure, operate, scale and govern their products, surfaced where they already work - the portal, the IDE (GitHub Copilot/MCP) and Teams. You will treat the platform as a product - shipping paved-road components other teams reuse, serving both humans and agents, with the multi-tenant scoping, cost-tracking, guardrails and governance an enterprise platform demands. RESPONSIBILITIES - Treat the platform as a product. Build paved roads and self-service: reusable AI building blocks (shared retrieval / "context engine," guardrail & evaluation libraries, an MCP/tool layer), scaffolder templates, SDK/API access and stable, versioned interfaces - built once, reused across features. - Ship AI experiences that delight developers. Grounded, well-cited assistants, copilots and wizards across the product lifecycle (e.g. a conversational knowledge assistant over our docs and catalog), meeting users on the portal, IDE (Copilot/MCP) and Teams via one shared API. - Serve humans and agents. Expose platform capabilities through an MCP / SDK / API surface - read-first, RBAC- and tenant-aware - so internal and external AI clients can query and (later, gated) act on the platform. See the AI-Assisted Platform Strategy RFC. - Own evaluation and quality. Build eval harnesses, golden tests and retrieval-quality metrics so features are correct, grounded and regression-tested in CI; invest in context engineering over model-shopping - the Gateway already solves model choice. - Pick the right pattern. Prefer deterministic pipelines + structured outputs + human-in-the-loop where outcomes are structured; reserve multi-step/multi-agent orchestration (Azure AI Foundry Agent Service, LangGraph / Microsoft Agent Framework) for genuinely open-ended tasks, keeping state-changing actions gated. - Strengthen MLOps / LLMOps. Improve prompt/version management, model adaptation, CI/CD
Requirements
and the path from experiment to production; treat prompts and retrieval as versioned, tested production assets. - Build for multi-tenancy. Default to per-product / per-tenant scoping of context, tools and actions; bake in observability (OpenTelemetry, Grafana, distributed tracing) and per-product cost/FinOps visibility. - Help advance security, safety & governance. Inherit platform RBAC (Entra ID / AccessIT), defend against the OWASP LLM Top 10, keep AI usage auditable, and respect BASF / EU AI Act and data-residency requirements. QUALIFICATIONS - BSc or MSc in Computer Science, Software Engineering, AI, or related field. - 4+ years in Software Engineering or Platform Development, with demonstrable recent experience in Generative AI and/or Agentic Systems. - You don't need to tick every box. Strong Python + hands-on LLM application experience + a platform/developer-experience mindset matter most; we expect you to grow into the rest. - AI / LLM engineering. Practical experience building LLM-powered applications; familiarity with RAG and agentic patterns (ReAct, plan-and-solve, multi-agent) and a clear sense of when not to use an autonomous agent. - Evaluation & quality (core). Designing eval harnesses, golden tests and retrieval-quality metrics for LLM/RAG systems (grounding, retrieval precision, hallucination control); context engineering over model selection. - Backend & API development. Python proficiency is highly desired (FastAPI, Pydantic, async); designing and operating production backend services and well-versioned APIs. - Platform / Developer-Experience engineering. Building reusable, self-service components and paved roads (templates, SDKs, golden paths) and operating multi-tenant services in pro