Researcher /x) in Visual Perception for Interactive 3D World
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+2 more
Job description
The core activities of the Department Augmented Vision at the German Research Center for Artificial Intelligence (DFKI) in Kaiserslautern lie in the fields of Image Processing and Computer Vision, Image Understanding, Augmented Reality, Virtual Reality, and 3D Reconstruction.
In the context of new national and European research projects, we are looking for highly motivated researchers to work on visual perception for interactive 3D worlds. Our research focuses on learning rich representations of humans, activities, and complex 3D environments using modern vision architectures, multimodal foundation models, world models, and end-to-end transformer-based approaches. Particular emphasis is placed on dynamic scene understanding, multimodal perception, vision-language reasoning, and embodied visual intelligence.
The position is intended for candidates who wish to pursue a PhD and contribute to high-quality scientific publications and research demonstrators.
Research and develop novel methods for visual perception and representation learning in interactive 3D environments.
Investigate multimodal perception, dynamic scene understanding, human activity understanding, world models, and vision-language reasoning.
Design and conduct rigorous experimental evaluations and contribute to scientific publications.
Implement research prototypes using modern vision architectures, transformer-based models, and GPU-based computing.
Contribute research components to demonstrators in Extended Reality, digital twins, human-AI interaction, and embodied AI.
Master’s degree or equivalent in Computer Science, Artificial Intelligence, Electrical Engineering, Robotics, Mathematics, or a related field.
Strong background in Computer Vision and machine learning.
Very good programming skills in Python and experience with modern research frameworks such as PyTorch.
Experience in some of the following areas: vision transformers, multimodal foundation models, world models, vision-language models, 3D Computer Vision, video understanding, human activity understanding, representation learning, or generative models.
Good understanding of modern neural architectures, large-scale representation learning, and experimental evaluation.
Strong motivation to conduct scientific research and pursue a PhD.
Excellent communication skills, ability to work independently, and willingness to contribute to a collaborative research environment.
Additional desirable qualifications
Experience with one or more of the following topics would be an advantage: end-to-end transformers, multimodal reasoning, video foundation models, Gaussian Splatting, neural scene representations, scene graphs, open-vocabulary perception, visual memory, continual learning, embodied AI, uncertainty-aware perception, event-based vision, or real-time visual systems.
Previous experience with scientific publications, open-source research code, benchmark datasets, or applied research projects is also welcome.
We offer excellent working conditions with challenging research topics in an interdisciplinary and international team at an internationally renowned research institute.
The position provides the opportunity to work on current research at the intersection of Computer Vision, multimodal perception, world models, 3D scene understanding, Extended Reality, and artificial intelligence. The successful candidate will contribute to visible research results, international publications, and demonstrators in collaboration with academic and industrial partners.
We expect the researcher to start or continue a PhD at RPTU Kaiserslautern-Landau during the project.
Requirements
Master’s degree or equivalent in Computer Science, Artificial Intelligence, Electrical Engineering, Robotics, Mathematics, or a related field.
Strong background in Computer Vision and machine learning.
Very good programming skills in Python and experience with modern research frameworks such as PyTorch.
Experience in some of the following areas: vision transformers, multimodal foundation models, world models, vision-language models, 3D Computer Vision, video understanding, human activity understanding, representation learning, or generative models.
Good understanding of modern neural architectures, large-scale representation learning, and experimental evaluation.
Strong motivation to conduct scientific research and pursue a PhD.
Excellent communication skills, ability to work independently, and willingness to contribute to a collaborative research environment.
Additional desirable qualifications
Experience with one or more of the following topics would be an advantage: end-to-end transformers, multimodal reasoning, video foundation models, Gaussian Splatting, neural scene representations, scene graphs, open-vocabulary perception, visual memory, continual learning, embodied AI, uncertainty-aware perception, event-based vision, or real-time visual systems.
Previous experience with scientific publications, open-source research code, benchmark datasets, or applied research projects is also welcome.
We offer excellent working conditions with challenging research topics in an interdisciplinary and international team at an internationally renowned research institute.
The position provides the opportunity to work on current research at the intersection of Computer Vision, multimodal perception, world models, 3D scene understanding, Extended Reality, and artificial intelligence. The successful candidate will contribute to visible research results, international publications, and demonstrators in collaboration with academic and industrial partners.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How to Find Tech Jobs in Berlin
The Most Popular IT Jobs on the Market
Dev Digest 121 - AI goes offline
System change: restart as developer?