AI Runtime / Low-Level Software Engineer

Kalray
Canton de Meylan, France
9 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Senior

Job location

Canton de Meylan, France

Tech stack

API
Artificial Intelligence
C++
Profiling
Software Quality
Computer Engineering
Concurrent Computing
Continuous Integration
Software Debugging
Linux
Embedded Software
Firmware
Linux kernel
Linux-Powered Devices
PCI Express
Performance Tuning
Quick EMUlator (QEMU)
Reduced Instruction Set Computing
Data Streaming
System Programming
Large Language Models
Information Technology
Bare Metal
Codebase
Hardware Acceleration

Job description

As an AI Runtime / Low-Level Software Engineer, you will own the offload runtime layer that enables inference workloads to run efficiently on our custom RISC-V AI accelerator from an x86 host.

You will join our AI & Compute team, which is building a full-stack GenAI inference platform, from serving to silicon: LLM Serving * AI Compilation * Runtime / Offload * Optimized AI Kernel Libraries

You will develop the Low-Level software stack responsible for device lifecycle management, scheduling and workload dispatch, high-performance host-to-device communication, and the runtime APIs exposed to compiler and serving layers. You will build this stack end-to-end on an open software foundation and collaborate closely with hardware, compiler and AI serving teams to drive the solution from emulation to production silicon.

Your main responsibilities will include:

Leading and contributing to:

  • Develop and optimize the runtime responsible for executing offloaded inference workloads, from the host offload API to the accelerator firmware.
  • Own the performance-critical path, including scheduling, dispatch, memory movement and data flow between the host and the accelerator.
  • Profile, analyze and optimize runtime performance by identifying bottlenecks related to latency, memory transfers and execution efficiency.
  • Design and maintain clean runtime APIs exposed to AI compiler and serving layers, enabling efficient workload execution on the accelerator.
  • Integrate and validate the real PCIe communication path while maintaining emulation environments as CI-tested development platforms.
  • Develop engineering foundations including CI/CD pipelines, automated smoke tests on emulated hardware and strong quality standards.
  • Collaborate with AI Inference, hardware and architecture teams to influence technical decisions and improve real-world AI performance.
  • Contribute to the continuous improvement and reliability of the complete AI acceleration software stack.

Requirements

  • Strong experience in C and/or C++ systems programming on Linux, within large and complex codebases.
  • Strong knowledge of parallel and concurrent programming, including threads, memory movement, profiling and performance optimization.
  • Experience with embedded software, firmware or bare-metal development, ideally with technologies such as OpenAMP, remoteproc, rpmsg, virtio and shared-memory communication models.
  • Good understanding of Linux internals, including concepts such as mmap, sysfs and the Linux device model, as well as PCIe fundamentals (BARs, DMA).
  • Experience with RISC-V (or similar ISA), cross-compilation and low-level debugging across the hardware/software boundary.

Nice to have:

  • Experience with the GenAI / LLM inference ecosystem, such as vLLM, or AI compilation technologies such as MLIR/LLVM.
  • Experience developing kernel drivers (PCIe, remoteproc) and working with QEMU or hardware emulation environments.
  • Experience with CI/CD systems targeting hardware or emulation platforms.
  • Experience with high-performance data-plane development, including DMA and zero-copy techniques, * MSc or Engineering degree (BAC+5) in Computer Science, Embedded Systems, Computer Engineering or related field.
  • 5+ years of experience in low-level software development, embedded systems, runtime development or systems programming.
  • Strong ownership mindset and autonomy with the ability to drive a complex software component end-to-end and make sound technical decisions.
  • Strong systems-thinking ability, understanding the complete chain from AI serving and compiler layers down to runtime, firmware and hardware.
  • Performance-oriented mindset with intuition for identifying where cycles, memory copies and bottlenecks impact execution.
  • Rigorous approach to software quality, automation, reproducibility and reliable engineering practices.
  • Strong collaboration and communication skills, with the ability to work closely with multidisciplinary teams.
  • Comfortable working in an international, fast-evolving startup environment.

Benefits & conditions

  • Competitive salary & performance-based RSU (free shares)
  • Hybrid work model
  • Additional paid leave (RTT)
  • Meal vouchers (Edenred)
  • Premium health coverage (Malakoff Humanis)
  • Sustainable mobility incentives
  • Generous paternity leave
  • Monthly team activities (laser game, hiking, sailing, karaoke …) and large-scale company events

About the company

Kalray is a European leader in hardware acceleration, with full-stack acceleration expertise: from silicon to complete system. Our MPPA® (Massively Parallel Processor Array) architecture is the foundation of Kalray's processor (30+ patent families, 15+ years of development) and acceleration cards that combine processing power, flexibility, and energy efficiency. Our mission is to deliver open data-efficient hardware accelerators to power next generation of data-intensive, AI-driven systems and infrastructures. We offer off-the-shelf processors, acceleration cards, and specialized processor development. With over 130 employees and presence in France and Romania, Kalray is backed by top-tier investors and publicly listed on Euronext Growth. You'll be part of a pioneering team that is shaping the future of computing with cutting-edge processor architecture, software-defined solutions, and next-generation acceleration platforms.

Apply for this position