System-Level Reliability Researcher

Leuven
Heverlee (Leuven), Belgium
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Abstraction Layers Data Analysis Architectural Patterns Computer Engineering Microarchitecture Emulators Fault Tolerance Systems Architecture Data Analytics Software Version Control Service Stack

Job description

  • Developing system-level reliability assessment methods for processors, accelerators, memory subsystems, SoCs, and chiplet-based architectures.
  • Creating architectural fault models and abstractions based on reliability information provided by lower layers of the technology stack.
  • Building and using simulation, emulation, and fault-injection frameworks to study fault propagation, error masking, and application-level impact.
  • Quantifying reliability outcomes such as Silent Data Corruption, detected errors, service interruptions, availability loss, and lifetime-related degradation at the system level.
  • Evaluating Reliability, Availability, and Serviceability (RAS) mechanisms, including error detection, containment, recovery, redundancy, and graceful degradation techniques.
  • Studying workload-dependent reliability behavior under realistic execution conditions and operating profiles.
  • Connecting technology-informed fault characteristics with system architecture models to support reliability-aware design decisions.
  • Collaborating closely with colleagues within and outside CSA to interface with workload models, architectural simulators, circuit-level reliability data, and technology-level observations.
  • Contributing to research publications, partner discussions, etc.

Requirements

  • You hold a Master’s or PhD in Electrical/Electronics/Computer Engineering/Science or relevant domains with strong R&D experience in system-level reliability.
  • You have 6+ years of relevant experience in industry and/or academia (we encourage you to apply if you have slightly less experience but strong alignment with the role).
  • You have a strong background in compute system architecture and microarchitecture.
  • You have a good understanding of core concepts in system-level reliability, resilience, and dependability, including:

*

  • Reliability, Availability, and Serviceability (RAS) architecture

*

  • Fault tolerance and resilience techniques

*

  • Error detection, containment, recovery, and graceful degradation

*

  • Fault propagation and architectural error masking

  • You have experience developing architectural models, quantitative analysis frameworks, or system-level evaluation methodologies.
  • You possess strong analytical and quantitative skills, including statistical analysis, data interpretation, and uncertainty-aware reasoning.
  • You have hands-on experience designing and executing experiments to derive insights from complex system behavior.
  • You are able to collaborate across disciplines and abstraction layers, translating reliability information from technology, circuit, and design teams into meaningful architectural models and system-level insights.
  • You have the knowledge and experience with good software and research engineering practices (e.g., version control, reproducibility, automation, testing, documentation).
  • You enjoy being hands-on while working with your colleagues and while mentoring students/interns, promoting a team culture of creativity, collaboration, and excellence.
  • You are a team player and flexible in accommodating changing priorities to support business needs; you are self-motivated, responsible and willing to take ownership.
  • You feel at home and know how to integrate into a multicultural environment; you are open-minded, you seek and embrace differences and accept constructive challenges.
  • You have effective English communication skills for discussions, documentations, and dissemination of ideas and results within CSA/IMEC and in wider ecosystems.

About the company

We offer you the opportunity to join one of the world’s premier research centers in nanotechnology at its headquarters in Leuven, Belgium. With your talent, passion and expertise, you’ll become part of a team that makes the impossible possible. Together, we shape the technology that will determine the society of tomorrow.

We are committed to being an inclusive employer and proud of our open, multicultural, and informal working environment with ample possibilities to take initiative and show responsibility. We commit to supporting and guiding you in this process; not only with words but also with tangible actions. Through imec.academy, ‘our corporate university’, we actively invest in your development to further your technical and personal growth.

We are aware that your valuable contribution makes imec a top player in its field. Your energy and commitment are therefore appreciated by means of a market appropriate salary with many fringe benefits.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.vdab.be

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:54 min

Applying reliability principles to software engineering leadership

Maxim Schepelin Maxim Schepelin · World Congress 2026 Europe

58 sec

Navigating limitations of local Cosmos DB emulators

Radu Vunvulea Radu Vunvulea · World Congress 2022

1:48 min

Automating exploratory data analysis within training pipelines

Dora Petrella · World Congress 2023

3:47 min

Solving the knowledge deficit in large language models

Alejandro Saucedo Alejandro Saucedo +3 · World Congress 2024

7:08 min

Engineering practices for extreme platform reliability

Justin Kitagawa · Coffee With Developers

1:49 min

Testing with emulators, simulators, and real devices

Milica Aleksic Milica Aleksic · LIVE

Videos

See all

Related articles

See all