Senior Software Engineer - Agent Safety / Evals - AI Foundations

Kraken
Berlin, Germany
10 days ago
Apply on de.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Cloud Computing Django Web Framework Python (Programming Language) Machine Learning Software Engineering Datadog Large Language Models Multi-Agent Systems Concurrency

Job description

You’ll work in the Agent Safety team. We build the shared platforms, harnesses, and guardrails that enable engineering and product teams to safely, reliably, and deterministically use machine learning and generative AI agents for internal systems and workflows across the business. This is a delivery-focused team that sits at the intersection of engineering, security, and internal enablement., We’re hiring a Senior Software Engineer to join our newly formed Agent Safety / Evals team. As AI agents take on more autonomous tasks across Kraken’s internal workflows, you will help ensure they do so securely, predictably, and within clearly defined operational boundaries.

This is a hands-on senior individual-contributor role. Working with the Lead Software Engineer and the broader AI team, you will own substantial parts of our evaluation, safety, and reliability stack - turning technical direction into production systems, contributing to architecture decisions, and raising engineering standards through thoughtful collaboration and mentoring.

What you’ll own

  • Build internal agent-safety systems: Design and implement services that improve the reliability, security, and determinism of LLMs and autonomous agents, keeping them within expected operational boundaries
  • Develop robust evaluation frameworks: Build scalable harnesses and evals for internal workflows and AI skills. Define meaningful test cases, interrogate output quality, and improve the reproducibility of the systems used to measure it
  • Implement guardrails and governance controls: Develop pre- and post-generation controls and translate agreed governance requirements into maintainable software across internal tools and platforms
  • Strengthen AI security and observability: Engineer authentication and permission patterns for internal agents, improve monitoring and tracing, and contribute to operational playbooks for AI-specific anomalies
  • Support verification and red teaming: Design tests and participate in continuous verification and red-teaming exercises to identify prompt injection, data-access, and unpredictable-behaviour risks before they reach production
  • Operate services in AWS: Deploy, run, and support high-throughput, low-latency safety services; make sound architecture, reliability, and cost trade-offs; and work effectively with platform, techops, and security partners
  • Raise the engineering bar: Contribute to technical decisions, review designs and code, mentor other engineers, and help establish pragmatic patterns that can be reused across AI Foundations

Requirements

  • Strong senior-level software engineering: Proven experience owning complex components or services end to end, from design and implementation through testing, deployment, and operation
  • Deep engineering fundamentals: Strong judgement around system design, concurrency, security, testing, and architecture trade-offs. Production Python experience is preferred
  • Practical AI evaluation and safety experience: A critical understanding of LLM behaviour and experience building or using evaluation harnesses to measure output quality and reliability
  • Security and governance mindset: Experience with areas such as threat modelling, authentication and authorisation, red teaming, data-access controls, or guardrails for internal platforms
  • Cloud experience: Comfortable running services in AWS, owning reliability and scalability, and collaborating with platform, techops, and security teams
  • Clear communication and collaboration: Able to explain technical trade-offs, challenge constructively, and turn complex safety and evaluation findings into practical action

What Success Looks Like

  • Reliable delivery: You ship well-tested safety and evaluation capabilities that are adopted by internal engineering and product teams
  • High-quality evals: Evaluation harnesses produce meaningful, reproducible signals that reflect the real-world performance and safety of Kraken’s internal AI tooling
  • Resilient infrastructure: The services you own are observable, secure, scalable, and supported by clear operational practices
  • Strong technical collaboration: You work effectively across AI Foundations and partner teams, improve designs through constructive challenge, and help others deliver safely and confidently

️ Bonus points

  • Evaluation and safety frameworks: Experience with tools such as Inspect AI, Ragas, OpenAI Evals, or NeMo Guardrails
  • Backend engineering: Django experience and strong patterns for security, performance, and maintainability
  • Observability: Experience using Datadog for tracing, monitoring, and investigating complex AI-enabled systems
  • AI engineering tooling: Familiarity with tools such as Pydantic AI, LiteLLM, or LangChain

About the company

Kraken powers some of the most innovative global developments in energy.

We create the technology that redefines utilities and unlocks a new energy system of the future. By optimising renewable generation, building a more intelligent grid, and empowering utilities to deliver an exceptional customer experience, our operating system is transforming the industry worldwide.

It’s an incredibly exciting time to work in energy. Join us on our mission to improve the lives of ONE BILLION humans within the decade and shape a cleaner, better future for everyone.

AI is a key investment area for Kraken Technologies as we look to expand our existing capabilities. A crucial part of this is broadening the foundational infrastructure that enables teams across the organisation to use AI effectively and accelerate our mission.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on de.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:36 min

Visualizing memory limits and isolating suspicious endpoints

Dina Matveev Dina Matveev ¡ Europe 2026 Virtual

5:59 min

Analyzing concurrency bottlenecks in standard serverless architectures

Marco Plaul Marco Plaul +1 ¡ World Congress 2023

2:27 min

Introduction to WebAssembly in a cloud computing context

Edo Edo ¡ World Congress 2024

3:31 min

Evolving developer roles into tech leads for AI agents

Alfonso Graziano Alfonso Graziano ¡ Coffee With Developers

1:08 min

Analyzing error logs and root causes using artificial intelligence

Nishil Patel Nishil Patel ¡ World Congress 2025

3:22 min

Evaluating advanced artificial intelligence platforms for daily recruitment

Rudi Bauer Rudi Bauer +1 ¡ Cappuccino with HR

Videos

See all

Related articles

See all