Sr. Embedded Machine Learning Engineer
Role details
Job location
Tech stack
Job description
We are looking for a Senior Embedded Machine Learning Engineer to own the end-to-end process of taking trained ML models and deploying them efficiently onto resource-constrained edge hardware. This role sits at the intersection of machine learning, embedded systems, and hardware engineering. You will integrate, convert, and optimize models to run within strict constraints on latency, memory, power, and thermal budget, and build the supporting C++ infrastructure that hosts them on device. You will partner closely with the CVML team who build the models, the embedded and firmware teams who own the device, and the product team who define performance targets. Success means models that are not just accurate in the lab but fast, small, and dependable in the field.
What You'll Do
- Apply quantization, pruning, knowledge distillation, operator fusion, and graph optimization to shrink models and reduce inference cost while protecting accuracy; convert trained models into edge-deployable formats using ONNX and TensorRT.
- Profile inference on target accelerators including GPUs, NPUs, DSPs, and FPGAs; measure latency, throughput, memory footprint, and power consumption, then drive the changes needed to hit performance targets.
- Design, write, and maintain the C++ application code that hosts inference on device, including pre- and post-processing pipelines, data and memory management, threading, and interfaces to the rest of the embedded system; ensure the combined model and C++ stack meets real-time constraints and fits within device memory budget.
- Build test harnesses to verify on-device accuracy against reference results and catch regressions from optimization or quantization; contribute to tooling for packaging, versioning, and delivering model updates to deployed devices.
- Set best practices for edge deployment, review designs and code, and mentor other engineers on optimization and embedded ML techniques; work closely with research, firmware, and product teams to set realistic performance targets and feed hardware constraints back into model design.
Requirements
- 10+ years of professional software or systems engineering experience, including at least 2 years focused on deploying ML models to embedded or edge devices; Bachelor's or Master's degree in Computer Science, Electrical Engineering, Computer Engineering, or equivalent practical experience.
- Very strong C++ proficiency; working knowledge of CUDA; hands-on experience with PyTorch and at least one edge inference runtime such as TensorFlow Lite, ONNX Runtime, or TensorRT.
- Practical experience with model optimization techniques including post-training quantization, quantization-aware training, pruning, and distillation; demonstrated ability to profile and optimize for latency, memory, and power on constrained hardware.
- Working knowledge of embedded or edge platforms such as NVIDIA Jetson, Qualcomm, ARM Cortex, or comparable NPUs and SoCs, and of Linux or an RTOS; solid grasp of computer architecture concepts relevant to inference including memory hierarchy, fixed-point arithmetic, and accelerator offload; domain experience in computer vision or sensor processing on device.
You'll Stand Out
- Hands-on experience deploying computer vision models for detection or tracking tasks on embedded or edge hardware.
- Experience with NVIDIA Jetson specifically, including TensorRT optimization and deployment on Jetson platforms.
- Background in defense, autonomous systems, or robotics where real-time reliability matters.
- Experience building or contributing to model update and OTA delivery pipelines for deployed edge devices.
Benefits & conditions
- Competitive salary
- ACS Equity Package
- Health, Dental, Vision Insurance
- Paid Time Off