Researcher /x) in Visual Perception for Interactive 3D World

DFKI GmbH
Germany
2 days ago
Apply on jobs.dfki.de
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
1 year minimum
Working hours
Regular working hours
Languages
German
Job source

Tech stack

Artificial Intelligence Computer Vision Augmented Reality Computer Programming Python (Programming Language) Machine Learning Language Modeling Open Source Technology Visual Systems Pytorch Virtual Reality Generative AI
+2 more
Gaussian Information Technology

Job description

The core activities of the Department Augmented Vision at the German Research Center for Artificial Intelligence (DFKI) in Kaiserslautern lie in the fields of Image Processing and Computer Vision, Image Understanding, Augmented Reality, Virtual Reality, and 3D Reconstruction.

In the context of new national and European research projects, we are looking for highly motivated researchers to work on visual perception for interactive 3D worlds. Our research focuses on learning rich representations of humans, activities, and complex 3D environments using modern vision architectures, multimodal foundation models, world models, and end-to-end transformer-based approaches. Particular emphasis is placed on dynamic scene understanding, multimodal perception, vision-language reasoning, and embodied visual intelligence.

The position is intended for candidates who wish to pursue a PhD and contribute to high-quality scientific publications and research demonstrators.

Research and develop novel methods for visual perception and representation learning in interactive 3D environments.

Investigate multimodal perception, dynamic scene understanding, human activity understanding, world models, and vision-language reasoning.

Design and conduct rigorous experimental evaluations and contribute to scientific publications.

Implement research prototypes using modern vision architectures, transformer-based models, and GPU-based computing.

Contribute research components to demonstrators in Extended Reality, digital twins, human-AI interaction, and embodied AI.

Master’s degree or equivalent in Computer Science, Artificial Intelligence, Electrical Engineering, Robotics, Mathematics, or a related field.

Strong background in Computer Vision and machine learning.

Very good programming skills in Python and experience with modern research frameworks such as PyTorch.

Experience in some of the following areas: vision transformers, multimodal foundation models, world models, vision-language models, 3D Computer Vision, video understanding, human activity understanding, representation learning, or generative models.

Good understanding of modern neural architectures, large-scale representation learning, and experimental evaluation.

Strong motivation to conduct scientific research and pursue a PhD.

Excellent communication skills, ability to work independently, and willingness to contribute to a collaborative research environment.

Additional desirable qualifications

Experience with one or more of the following topics would be an advantage: end-to-end transformers, multimodal reasoning, video foundation models, Gaussian Splatting, neural scene representations, scene graphs, open-vocabulary perception, visual memory, continual learning, embodied AI, uncertainty-aware perception, event-based vision, or real-time visual systems.

Previous experience with scientific publications, open-source research code, benchmark datasets, or applied research projects is also welcome.

We offer excellent working conditions with challenging research topics in an interdisciplinary and international team at an internationally renowned research institute.

The position provides the opportunity to work on current research at the intersection of Computer Vision, multimodal perception, world models, 3D scene understanding, Extended Reality, and artificial intelligence. The successful candidate will contribute to visible research results, international publications, and demonstrators in collaboration with academic and industrial partners.

We expect the researcher to start or continue a PhD at RPTU Kaiserslautern-Landau during the project.

Requirements

Master’s degree or equivalent in Computer Science, Artificial Intelligence, Electrical Engineering, Robotics, Mathematics, or a related field.

Strong background in Computer Vision and machine learning.

Very good programming skills in Python and experience with modern research frameworks such as PyTorch.

Experience in some of the following areas: vision transformers, multimodal foundation models, world models, vision-language models, 3D Computer Vision, video understanding, human activity understanding, representation learning, or generative models.

Good understanding of modern neural architectures, large-scale representation learning, and experimental evaluation.

Strong motivation to conduct scientific research and pursue a PhD.

Excellent communication skills, ability to work independently, and willingness to contribute to a collaborative research environment.

Additional desirable qualifications

Experience with one or more of the following topics would be an advantage: end-to-end transformers, multimodal reasoning, video foundation models, Gaussian Splatting, neural scene representations, scene graphs, open-vocabulary perception, visual memory, continual learning, embodied AI, uncertainty-aware perception, event-based vision, or real-time visual systems.

Previous experience with scientific publications, open-source research code, benchmark datasets, or applied research projects is also welcome.

We offer excellent working conditions with challenging research topics in an interdisciplinary and international team at an internationally renowned research institute.

The position provides the opportunity to work on current research at the intersection of Computer Vision, multimodal perception, world models, 3D scene understanding, Extended Reality, and artificial intelligence. The successful candidate will contribute to visible research results, international publications, and demonstrators in collaboration with academic and industrial partners.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on jobs.dfki.de
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:04 min

Academic research in games and virtual reality

Chris Heilmann +2 · LIVE

2:37 min

Tracing the evolution from early AI to generative AI

Mike Mike · World Congress 2025

1:43 min

Mitigating vision classifier attacks using Gaussian blur techniques

David vonThenen David vonThenen · World Congress 2025

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

2:14 min

Introduction to generative AI and content warnings

Cheuk Ho · World Congress 2023

2:21 min

Applying diffusion models for image upscaling and refinement

Han Xiao · World Congress 2022

Videos

See all

Related articles

See all