Materials Data Engineer (AI for Science, Knowledge Graphs) - Developer

Llms
Karlsruhe, Germany
3 days ago
Apply on www.careerjet.de
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Shift work

Tech stack

JavaScript (Programming Language) Artificial Intelligence Databases Graph Database Information Sciences Python (Programming Language) Machine Learning Web Ontology Language Open Source Technology Resource Description Framework (RDF) Tensorflow Software Engineering
+9 more
SPARQL AI Infrastructure Digital Twin Pytorch Large Language Models Build Management Information Technology Production Code Virtual Agents

Job description

Traditional scientific knowledge is still largely hidden from LLMs because physical R&D data (like lab experiments, simulations, and equipment logs) is rarely recorded in a structured, machine-actionable way. Text alone isn’t rich enough to support automated discovery. To bridge this gap, we have built an ontology-driven, schema-based knowledge graph management system. Now, we are taking it to the next level: building autonomous, goal-oriented AI agents that can interact directly with our graph databases, augment them with new data, and identify emerging patterns in physical science. This role offers a unique opportunity to design production-grade AI agent systems from scratch, collaborating closely with experienced material scientists, tribologists, and software engineers. At datin, we value curiosity, impact, and trust, and we design our agent-driven workflows to empower scientists, not replace them. Tasks

  • Agentic Workflows: Design and build end-to-end agentic architectures. You will build tool-calling loops, memory layers, and execution environments that allow agents to query, update, and validate our graph databases.
  • AI Infrastructure: Engineer, deploy, and maintain performant agent and LLM serving infrastructures both locally and in the cloud.
  • Graph-Grounded LLMs: Fine-tune or optimize open-source LLMs to reliably translate natural language scientific requests into structured queries sent to our SDK and accurately traverse complex ontologies.
  • Machine Learning for Science: Train and integrate specialized ML models to solve multi-objective optimization problems (e.g., predicting material properties or chemical reactions) that AI agents can use as tools.
  • Semantic Digital Twins: Translate real-world physical workflows into semantically-typed knowledge graphs.

Requirements

  • Technical Core: Deep practical experience with Agentic frameworks, orchestrators, or tool-use libraries.
  • Software Engineering: Strong proficiency in Python and/or JavaScript, with a focus on writing clean, modular, and well-tested production code.
  • Modeling Skills: Hands-on experience building, training, or fine-tuning models using machine learning frameworks like PyTorch or similar.
  • Validation: Familiarity with SHACL, RDF, RDFS, OWL, and SPARQL or similar (like CYPHER) validation languages is a strong plus.
  • Background: A degree in Computer Science, Information Science, or, Chemistry, Materials Science, Mechanical Engineering, or a related field.
  • Mindset: You are meticulous and logical. You enjoy solving the puzzle of how to structure the world into a database.

Benefits & conditions

  • Flexible working hours
  • Free beverages
  • Public transportation benefits
  • Remote work possible
  • Travel expenses compensation

About the company

This is the future of scientific AI. At datin GmbH, you won’t just be writing code; you will be defining the grammar of scientific discovery. If you are ready to build the engine that powers the next generation of R&D, apply now and let’s shape the future together. datin GmbH Scientific discovery is the engine of human progress, yet the tools researchers use today are stuck in the past: paper notebooks, scattered PDFs, and rigid spreadsheets. datin is here to change that. Based in Karlsruhe, we are building the world’s first AI-native infrastructure for Research & Development. We replace complex data silos with an intuitive Mind-Map interface that mirrors how scientists actually think. Behind the scenes, our platform weaves these workflows into a powerful Knowledge Graph, creating a living “corporate memory.” This allows R&D teams to move from messy trial-and-error to true AI-assisted development, ensuring that no experimental data is ever lost or wasted. We are a startup, founded by experienced scientists. We don’t just plug into the wall; we are building the new power grid for innovation. From materials science to green energy, our software enables the breakthroughs that shape our future.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.de
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:04 min

Database evolution and the funding behind vector databases

Erik Bamberg · LIVE

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

4:01 min

Managing application isolation via pluggable database models

Wei Hu Wei Hu · World Congress 2022

1:06 min

Compiling PyTorch environments for advanced time forecasting

Christoph Lohrmann Christoph Lohrmann +1 · World Congress 2026 Europe

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

Videos

See all

Related articles

See all