Sr. Embedded Machine Learning Engineer

Allen Control Systems
Mountain View, United States of America
yesterday

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Senior
Compensation
$ 261K

Job location

Mountain View, United States of America

Tech stack

Systems Engineering
ARM
Computer Vision
C++
Program Optimization
Nvidia CUDA
Computer Engineering
Linux
Memory Management
Firmware
Field-Programmable Gate Array (FPGA)
Machine Learning
Real-Time Operating Systems
TensorFlow
Smart Devices
Management of Software Versions
Graphics Processing Unit (GPU)
PyTorch
Information Technology
Low Latency
ONNX (Open Neural Network Exchange) Format
Machine Learning Operations
TensorRT
Data Pipelines

Job description

We are looking for a Senior Embedded Machine Learning Engineer to own the end-to-end process of taking trained ML models and deploying them efficiently onto resource-constrained edge hardware. This role sits at the intersection of machine learning, embedded systems, and hardware engineering. You will integrate, convert, and optimize models to run within strict constraints on latency, memory, power, and thermal budget, and build the supporting C++ infrastructure that hosts them on device. You will partner closely with the CVML team who build the models, the embedded and firmware teams who own the device, and the product team who define performance targets. Success means models that are not just accurate in the lab but fast, small, and dependable in the field.

What You'll Do

  • Apply quantization, pruning, knowledge distillation, operator fusion, and graph optimization to shrink models and reduce inference cost while protecting accuracy; convert trained models into edge-deployable formats using ONNX and TensorRT.
  • Profile inference on target accelerators including GPUs, NPUs, DSPs, and FPGAs; measure latency, throughput, memory footprint, and power consumption, then drive the changes needed to hit performance targets.
  • Design, write, and maintain the C++ application code that hosts inference on device, including pre- and post-processing pipelines, data and memory management, threading, and interfaces to the rest of the embedded system; ensure the combined model and C++ stack meets real-time constraints and fits within device memory budget.
  • Build test harnesses to verify on-device accuracy against reference results and catch regressions from optimization or quantization; contribute to tooling for packaging, versioning, and delivering model updates to deployed devices.
  • Set best practices for edge deployment, review designs and code, and mentor other engineers on optimization and embedded ML techniques; work closely with research, firmware, and product teams to set realistic performance targets and feed hardware constraints back into model design.

Requirements

  • 10+ years of professional software or systems engineering experience, including at least 2 years focused on deploying ML models to embedded or edge devices; Bachelor's or Master's degree in Computer Science, Electrical Engineering, Computer Engineering, or equivalent practical experience.
  • Very strong C++ proficiency; working knowledge of CUDA; hands-on experience with PyTorch and at least one edge inference runtime such as TensorFlow Lite, ONNX Runtime, or TensorRT.
  • Practical experience with model optimization techniques including post-training quantization, quantization-aware training, pruning, and distillation; demonstrated ability to profile and optimize for latency, memory, and power on constrained hardware.
  • Working knowledge of embedded or edge platforms such as NVIDIA Jetson, Qualcomm, ARM Cortex, or comparable NPUs and SoCs, and of Linux or an RTOS; solid grasp of computer architecture concepts relevant to inference including memory hierarchy, fixed-point arithmetic, and accelerator offload; domain experience in computer vision or sensor processing on device.

You'll Stand Out

  • Hands-on experience deploying computer vision models for detection or tracking tasks on embedded or edge hardware.
  • Experience with NVIDIA Jetson specifically, including TensorRT optimization and deployment on Jetson platforms.
  • Background in defense, autonomous systems, or robotics where real-time reliability matters.
  • Experience building or contributing to model update and OTA delivery pipelines for deployed edge devices.

Benefits & conditions

  • Competitive salary
  • ACS Equity Package
  • Health, Dental, Vision Insurance
  • Paid Time Off

About the company

Allen Control Systems (ACS) is a cutting-edge defense startup founded by two former Navy electrical engineers with a proven track record in robotics and software. We are developing an autonomous gun turret using advanced computer vision and control systems to precisely detect, track, and neutralize enemy drones. With an engineering-first culture, ACS values technical excellence and innovation. Backed by our founders' successful exits from two previous ventures acquired for a combined $180M in 2022, we are committed to ensuring that the groundbreaking technologies we develop will have a real-world impact.

Apply for this position