[Remote in US] AI Kernel Engineer - RISC-V Software Stack
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+25 more
Job description
You will work at the intersection of embedded systems, computer architecture, and machine learning, developing high-performance compute kernels, runtime components, libraries, and developer-facing SDK tools. A key part of the role will be efficiently mapping compute-intensive workloads such as convolution, matrix multiplication, and signal-processing operations onto a multicore RISC-V SoC.
The role focuses heavily on vector/SIMD execution, memory optimization, data movement, multicore parallelism, and low-level performance optimization.
Prior RISC-V experience is valuable but not required. Engineers with backgrounds in ARM NEON/SVE, x86 SIMD/AVX, DSP software, GPU kernel programming, embedded performance optimization, or other low-level parallel architectures are encouraged to apply.
You will collaborate closely with RTL design, system architecture, software, and machine learning teams to turn architectural capabilities into a practical, high-performance, and extensible software platform.
Key Responsibilities
- Develop and optimize high-performance ML and DSP compute kernels, including operations such as convolution, matrix multiplication, activation functions, pooling, image-processing primitives, and related numerical workloads
- Optimize computationally intensive C/C++ code for vector/SIMD execution, multicore processing, and the SoC memory hierarchy
- Build reusable compute libraries, runtime components, APIs, and developer-facing components for the Mentium SDK
- Develop efficient data-movement, memory-management, and workload-scheduling strategies
- Optimize the use of caches, scratchpad memories, DMA engines, and on-chip memory resources
- Profile workloads and identify compute, memory-bandwidth, synchronization, and system-level performance bottlenecks
- Perform low-level performance analysis using profiling, benchmarking, cycle measurements, and hardware/software debugging tools
- Integrate optimized compute kernels and runtime components into AI model deployment and inference workflows
- Develop functional tests, performance benchmarks, reference examples, and SDK documentation
- Collaborate closely with RTL and system-architecture engineers to validate hardware features and improve end-to-end system performance
- Contribute to the architecture and programming model of Mentium’s RISC-V accelerator software stack
- Evaluate and adapt relevant open-source libraries, runtimes, compiler technologies, and numerical software
- Help define software requirements and provide feedback that influences future hardware architecture, * Quantized inference, fixed-point arithmetic, INT8/INT16 computation, FP16/BF16, or other reduced-precision numerical formats
- DMA, scratchpad memory, cache hierarchies, memory bandwidth optimization, and multicore synchronization
- Embedded, bare-metal, real-time, or resource-constrained software development
- Multicore SoCs or heterogeneous compute architectures
- Open-source RISC-V platforms such as PULP or similar multicore/accelerator systems
- Machine-learning frameworks and model formats such as PyTorch, TensorFlow, TFLite, or ONNX
- Compiler and deployment technologies such as LLVM, MLIR, TVM, Deeploy, or related systems
- SDKs, runtime libraries, numerical libraries, developer tools, or reusable software APIs
- Hardware-software co-design, SoC development, FPGA prototyping, architectural simulation, or custom accelerator development
- Open-source software or research software development
Why Join Mentium?
At Mentium, you will work at the intersection of custom silicon, RISC-V, high-performance embedded software, and AI.
You will work directly with the engineers designing the underlying hardware and play a central role in determining how developers and machine-learning workloads interact with our accelerator.
Rather than simply programming an existing processor, you will have the opportunity to influence the hardware-software boundary: identifying architectural bottlenecks, developing optimized compute kernels, evaluating new programming approaches, and providing feedback that can shape future generations of the hardware.
Requirements
- Bachelor’s degree in Electrical Engineering, Computer Engineering, Computer Science, or a related technical field, or equivalent practical experience
- 3+ years of combined relevant industry, graduate research, doctoral research, or applied research experience
- Strong programming skills in C and/or C++
- Experience developing or optimizing performance-critical software
- Experience with at least one area of low-level performance programming, such as:
- SIMD or vector programming
- DSP programming
- GPU kernel programming
- Assembly or intrinsic-based optimization
- Performance-critical embedded software
- Numerical or high-performance computing
- Solid understanding of computer architecture, memory systems, and parallel processing
- Experience with performance profiling, benchmarking, low-level debugging, or cycle-level optimization
- Familiarity with computational workloads such as convolution, matrix multiplication, image processing, signal processing, or other numerical kernels
- Ability to reason about memory access patterns, data locality, computational efficiency, and hardware utilization
- Ability to read hardware specifications and work effectively with hardware and RTL engineers
- Proficiency with Python for testing, automation, benchmarking, tooling, or application development
- Experience with Git and standard collaborative software-development practices
- Strong written and verbal communication skills, Experience in several of the following areas is valuable, but we do not expect candidates to have experience with all of them:
- RISC-V instruction-set architecture or the RISC-V Vector Extension (RVV)
- ARM NEON or SVE, x86 SSE/AVX, DSP vector architectures, GPUs, or other SIMD/vector processors
- Vector intrinsics, assembly programming, compiler intrinsics, or low-level code optimization
Benefits & conditions
- Competitive compensation packages
- Opportunity to work on diverse, cutting-edge AI projects across a range of industries.
- 401(k)
- Flexible PTO
- Full PPO medical, dental, and vision insurance coverage
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Dev Digest 120 - Apple and peers
Highest Paying Tech Companies for Developers
Is Software Engineering Over-Saturated?
Fully Remote Software Engineer Jobs