NPU Architect

SEMRON
Dresden, Germany
about 1 month ago
Apply on de.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence C++ (Programming Language) Computer Programming Microarchitecture Logic Synthesis of Circuits Hardware Description Language Python (Programming Language) Machine Learning Software Deployment SystemVerilog Verilog

Job description

In this role you’ll be responsible to design the next iteration of SEMRON’s 3D in-memory compute chip. You will collaborate with a team of hardware, compiler and ML engineers to optimise all aspects of executing ML workloads based on our capacitive analog in-memory matrix-vector multiplication units.

What you will do:

  • Design and specify the structure and internal organisation of core architecture modules, aligned with workload requirements and software deployment processes.
  • Partner with the software team to evaluate module performance and efficiency for key workloads, uncover performance constraints, and inform architectural choices.
  • Monitor emerging trends and research in AI workloads, hardware architectures, and applications to guide the evolution of next-generation architectures.

Requirements

  • BS/MS/PhD in EE, CS, or a related field
  • Understanding in computer architecture, digital design, and micro-architecture concepts
  • Familiarity with AI/ML algorithms, frameworks, and workloads
  • Programming experience in C/C++ and Python, * Hands-On Experience with ML Hardware Exploration Frameworks like Timeloop, ZigZag, etc.
  • Hands-On Experience with HDLs such as Verilog or System Verilog

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on de.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:36 min

Applying supervised machine learning for practical rule extraction

Katja Träumner

6:59 min

Deep dive into nemotron ultra architecture and performance

Sergio Perez Sergio Perez · World Congress 2026 Europe

1:57 min

Evolution of machine learning algorithms and computing hardware

Alexandra Waldherr · LIVE

3:30 min

Transitioning from CUDA software architect to user

Stephen Jones · Coffee With Developers

1:28 min

Exploring quantum execution frameworks and hardware layouts

Alex Waldherr Alex Waldherr · World Congress 2022

3:51 min

Overcoming hardware configuration barriers in machine learning

Jose Luis Latorre Millas · LIVE

Videos

See all

Related articles

See all