AI Runtime / Low-Level Software Engineer

Kalray
Montbonnot-Saint-Martin, France
21 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence C++ (Programming Language) Profiling Software Quality Computer Engineering Concurrent Computing Continuous Integration Software Debugging Linux Embedded Software Firmware
+13 more
Linux Kernel Linux-Powered Devices PCI Express Performance Tuning Quick EMUlator (QEMU) Reduced Instruction Set Computing Data Streaming System Programming Large Language Models Information Technology Bare Metal Codebase Hardware Acceleration

Job description

As an AI Runtime / Low-Level Software Engineer, you will own the offload runtime layer that enables inference workloads to run efficiently on our custom RISC-V AI accelerator from an x86 host.

You will join our AI & Compute team, which is building a full-stack GenAI inference platform, from serving to silicon: LLM Serving * AI Compilation * Runtime / Offload * Optimized AI Kernel Libraries

You will develop the Low-Level software stack responsible for device lifecycle management, scheduling and workload dispatch, high-performance host-to-device communication, and the runtime APIs exposed to compiler and serving layers. You will build this stack end-to-end on an open software foundation and collaborate closely with hardware, compiler and AI serving teams to drive the solution from emulation to production silicon.

Your main responsibilities will include:

Leading and contributing to:

  • Develop and optimize the runtime responsible for executing offloaded inference workloads, from the host offload API to the accelerator firmware.
  • Own the performance-critical path, including scheduling, dispatch, memory movement and data flow between the host and the accelerator.
  • Profile, analyze and optimize runtime performance by identifying bottlenecks related to latency, memory transfers and execution efficiency.
  • Design and maintain clean runtime APIs exposed to AI compiler and serving layers, enabling efficient workload execution on the accelerator.
  • Integrate and validate the real PCIe communication path while maintaining emulation environments as CI-tested development platforms.
  • Develop engineering foundations including CI/CD pipelines, automated smoke tests on emulated hardware and strong quality standards.
  • Collaborate with AI Inference, hardware and architecture teams to influence technical decisions and improve real-world AI performance.
  • Contribute to the continuous improvement and reliability of the complete AI acceleration software stack.

Requirements

  • Strong experience in C and/or C++ systems programming on Linux, within large and complex codebases.
  • Strong knowledge of parallel and concurrent programming, including threads, memory movement, profiling and performance optimization.
  • Experience with embedded software, firmware or bare-metal development, ideally with technologies such as OpenAMP, remoteproc, rpmsg, virtio and shared-memory communication models.
  • Good understanding of Linux internals, including concepts such as mmap, sysfs and the Linux device model, as well as PCIe fundamentals (BARs, DMA).
  • Experience with RISC-V (or similar ISA), cross-compilation and low-level debugging across the hardware/software boundary.

Nice to have:

  • Experience with the GenAI / LLM inference ecosystem, such as vLLM, or AI compilation technologies such as MLIR/LLVM.
  • Experience developing kernel drivers (PCIe, remoteproc) and working with QEMU or hardware emulation environments.
  • Experience with CI/CD systems targeting hardware or emulation platforms.
  • Experience with high-performance data-plane development, including DMA and zero-copy techniques, * MSc or Engineering degree (BAC+5) in Computer Science, Embedded Systems, Computer Engineering or related field.
  • 5+ years of experience in low-level software development, embedded systems, runtime development or systems programming.
  • Strong ownership mindset and autonomy with the ability to drive a complex software component end-to-end and make sound technical decisions.
  • Strong systems-thinking ability, understanding the complete chain from AI serving and compiler layers down to runtime, firmware and hardware.
  • Performance-oriented mindset with intuition for identifying where cycles, memory copies and bottlenecks impact execution.
  • Rigorous approach to software quality, automation, reproducibility and reliable engineering practices.
  • Strong collaboration and communication skills, with the ability to work closely with multidisciplinary teams.
  • Comfortable working in an international, fast-evolving startup environment.

Benefits & conditions

  • Competitive salary & performance-based RSU (free shares)
  • Hybrid work model
  • Additional paid leave (RTT)
  • Meal vouchers (Edenred)
  • Premium health coverage (Malakoff Humanis)
  • Sustainable mobility incentives
  • Generous paternity leave
  • Monthly team activities (laser game, hiking, sailing, karaoke …) and large-scale company events

About the company

Kalray is a European leader in hardware acceleration, with full-stack acceleration expertise: from silicon to complete system.

Our MPPA® (Massively Parallel Processor Array) architecture is the foundation of Kalray’s processor (30+ patent families, 15+ years of development) and acceleration cards that combine processing power, flexibility, and energy efficiency.

Our mission is to deliver open data-efficient hardware accelerators to power next generation of data-intensive, AI-driven systems and infrastructures. We offer off-the-shelf processors, acceleration cards, and specialized processor development.

With over 130 employees and presence in France and Romania, Kalray is backed by top-tier investors and publicly listed on Euronext Growth. You’ll be part of a pioneering team that is shaping the future of computing with cutting-edge processor architecture, software-defined solutions, and next-generation acceleration platforms.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.kalrayinc.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

3:34 min

Augmenting codebases with semantic architectures

Zaak Chalal Zaak Chalal · World Congress 2026 Europe

2:19 min

Orchestrating over-the-air firmware updates for vehicle modules

Denis Grahovac · World Congress 2021

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

2:37 min

Optimizing technical profiles for AI sourcing and recruitment

Mina Golesorkhi Mina Golesorkhi · World Congress 2026 Europe

Videos

See all

Related articles

See all