> Markdown version of [/jobs/ext/2492967-senior-software-engineer-agent-safety-evals-ai-foundations](https://www.wearedevelopers.com/jobs/ext/2492967-senior-software-engineer-agent-safety-evals-ai-foundations). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Software Engineer - Agent Safety / Evals - AI Foundations - **Company:** Kraken - **Location:** Berlin, Germany - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Cloud Computing, Django Web Framework, Python (Programming Language), Machine Learning, Software Engineering, Datadog, Large Language Models, Multi-Agent Systems, Concurrency - **Published:** August 25, 2026 - **Apply:** https://de.indeed.com/viewjob?jk=76f6e4efde9dbce9 ## About the Role * Strong senior-level software engineering: Proven experience owning complex components or services end to end, from design and implementation through testing, deployment, and operation * Deep engineering fundamentals: Strong judgement around system design, concurrency, security, testing, and architecture trade-offs. Production Python experience is preferred * Practical AI evaluation and safety experience: A critical understanding of LLM behaviour and experience building or using evaluation harnesses to measure output quality and reliability * Security and governance mindset: Experience with areas such as threat modelling, authentication and authorisation, red teaming, data-access controls, or guardrails for internal platforms * Cloud experience: Comfortable running services in AWS, owning reliability and scalability, and collaborating with platform, techops, and security teams * Clear communication and collaboration: Able to explain technical trade-offs, challenge constructively, and turn complex safety and evaluation findings into practical action What Success Looks Like * Reliable delivery: You ship well-tested safety and evaluation capabilities that are adopted by internal engineering and product teams * High-quality evals: Evaluation harnesses produce meaningful, reproducible signals that reflect the real-world performance and safety of Kraken's internal AI tooling * Resilient infrastructure: The services you own are observable, secure, scalable, and supported by clear operational practices * Strong technical collaboration: You work effectively across AI Foundations and partner teams, improve designs through constructive challenge, and help others deliver safely and confidently ️ Bonus points * Evaluation and safety frameworks: Experience with tools such as Inspect AI, Ragas, OpenAI Evals, or NeMo Guardrails * Backend engineering: Django experience and strong patterns for security, performance, and maintainability * Observability: Experience using Datadog for tracing, monitoring, and investigating complex AI-enabled systems * AI engineering tooling: Familiarity with tools such as Pydantic AI, LiteLLM, or LangChain ## Description You'll work in the Agent Safety team. We build the shared platforms, harnesses, and guardrails that enable engineering and product teams to safely, reliably, and deterministically use machine learning and generative AI agents for internal systems and workflows across the business. This is a delivery-focused team that sits at the intersection of engineering, security, and internal enablement., We're hiring a Senior Software Engineer to join our newly formed Agent Safety / Evals team. As AI agents take on more autonomous tasks across Kraken's internal workflows, you will help ensure they do so securely, predictably, and within clearly defined operational boundaries. This is a hands-on senior individual-contributor role. Working with the Lead Software Engineer and the broader AI team, you will own substantial parts of our evaluation, safety, and reliability stack - turning technical direction into production systems, contributing to architecture decisions, and raising engineering standards through thoughtful collaboration and mentoring. What you'll own * Build internal agent-safety systems: Design and implement services that improve the reliability, security, and determinism of LLMs and autonomous agents, keeping them within expected operational boundaries * Develop robust evaluation frameworks: Build scalable harnesses and evals for internal workflows and AI skills. Define meaningful test cases, interrogate output quality, and improve the reproducibility of the systems used to measure it * Implement guardrails and governance controls: Develop pre- and post-generation controls and translate agreed governance requirements into maintainable software across internal tools and platforms * Strengthen AI security and observability: Engineer authentication and permission patterns for internal agents, improve monitoring and tracing, and contribute to operational playbooks for AI-specific anomalies * Support verification and red teaming: Design tests and participate in continuous verification and red-teaming exercises to identify prompt injection, data-access, and unpredictable-behaviour risks before they reach production * Operate services in AWS: Deploy, run, and support high-throughput, low-latency safety services; make sound architecture, reliability, and cost trade-offs; and work effectively with platform, techops, and security partners * Raise the engineering bar: Contribute to technical decisions, review designs and code, mentor other engineers, and help establish pragmatic patterns that can be reused across AI Foundations ## Related Videos - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [Concurrency with Go](https://www.wearedevelopers.com/videos/191-concurrency-with-go) - [Green Cloud Computing](https://www.wearedevelopers.com/videos/592-green-cloud-computing) - [Agentic AI: building autonomous systems for developers](https://www.wearedevelopers.com/videos/100034-agentic-ai-building-autonomous-systems-for-developers) - [WebAssembly: The Next Frontier of Cloud Computing](https://www.wearedevelopers.com/videos/972-webassembly-the-next-frontier-of-cloud-computing) - [Structured Concurrency in Practice: CoroutineScope vs StructuredTaskScope](https://www.wearedevelopers.com/videos/100349-structured-concurrency-in-practice-coroutinescope-vs-structuredtaskscope) ## Related Articles - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [What is Software Engineering in the Age of AI?](https://www.wearedevelopers.com/magazine/640-what-is-software-engineering-in-the-age-of-ai) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)