ML Systems Engineer - Inference Acceleration

Arago
Courbevoie, France
1 day ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Job source

Tech stack

Computer-Aided Design C++ (Programming Language) Nvidia CUDA Computer Programming Extract Transform Load (ETL) Python (Programming Language) System Programming Large Language Models Caching Parallel Computation Machine Learning Operations TensorRT

Job description

  • Analyze modern AI workloads and identify kernel-, runtime-, memory-, and system-level bottlenecks on Arago’s accelerator.
  • Develop and optimize custom kernels, fused operators, and execution strategies to maximize device utilization.
  • Design efficient mappings of models and operators across multiple Arago devices, including communication and synchronization strategies.
  • Develop inference-serving techniques such as continuous batching, paged KV caches, prefix/context caching, chunked prefill, and prefill/decode interleaving or disaggregation.
  • Build profiling, benchmarking, and performance-analysis infrastructure spanning kernels, full models, and serving workloads.
  • Work closely with Arago’s hardware, compiler, and runtime teams to co-design software abstractions and influence future hardware features based on real model workloads.

Requirements

  • Strong experience in high-performance ML inference, GPU/accelerator programming, or ML systems engineering.
  • Deep understanding of computer architecture, accelerator/GPU execution models, memory hierarchies, parallelism, and performance bottlenecks.
  • Experience developing and optimizing custom kernels using CUDA, Triton, ROCm/HIP, or equivalent low-level programming environments.
  • Experience with operator fusion, tiling, scheduling, data movement optimization, graph execution, and profiling of compute- and memory-bound workloads.
  • Strong understanding of distributed model execution, including tensor, pipeline, sequence, and/or expert parallelism and communication/computation overlap.
  • Hands-on experience with modern inference-serving systems such as vLLM, SGLang, TensorRT-LLM, or equivalent, including KV-cache management, continuous batching, paged attention, and prefill/decode scheduling.
  • Strong C++ and Python skills, and comfort working on a custom accelerator stack where compiler, runtime, kernels, and abstractions are actively being developed. Exposure to or experience with MLIR and MLIR dialects is a strong plus.
  • Language: English at a proficient level.

Benefits & conditions

  • Competitive cash compensation, with final package based on location, experience, and the pay of team members in similar positions.
  • Meaningful stock option plan offered (included in the majority of full time offers).
  • Healthcare coverage (including family-friendly options), pension contributions, professional development support, and 25 days of PTO, in addition to public holidays.
  • Ownership of a key technical domain, with significant vertical and/or horizontal growth opportunities, based on performance and individual drive.

About the company

The explosive growth of AI is pushing the industry to rethink how processors are built. Arago is meeting that challenge with a proprietary technology that fuses optical and CMOS technologies to deliver an order-of-magnitude increase in performance.

Arago is the fastest, and currently the only, company to have built such a processor. It’s backed by leading deep-tech investors and some of the most respected figures in semiconductors and computing, including the CEO of Arm, the founder of macOS who worked directly with Steve Jobs at Apple, an Nvidia Fellow, the Head of Optics at Google, and many other industry leaders.

Our work is guided by three clear values: do great things, move with high velocity, and operate as one unit. We work in a demanding environment where constant learning, ownership, and execution are expected, and where exceptional people have the opportunity to do their life’s work.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:32 min

Core libraries driving inference engines and multi-GPU networking

Adolf Hohl Adolf Hohl · World Congress 2024

6:21 min

Previewing upcoming hardware acceleration capabilities for Python environments

Chris Heilmann +2 · LIVE

4:52 min

Essential phases in building and refining language models

Anshul Jindal Anshul Jindal +1 · World Congress 2025

3:15 min

Reversing the caching model for artifact delivery

Thijs Feryn Thijs Feryn · World Congress 2026 Europe

1:37 min

Accelerating compute with focused developer tools

Julia Koch Julia Koch +1 · World Congress 2026 Europe

2:33 min

Maintaining prompt structures for prefix caching

Douglas Reiser Douglas Reiser · Europe 2026 Virtual

Videos

See all

Related articles

See all