> Markdown version of [/jobs/ext/1920842-sr-embedded-machine-learning-engineer](https://www.wearedevelopers.com/jobs/ext/1920842-sr-embedded-machine-learning-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Sr. Embedded Machine Learning Engineer - **Company:** Allen Control Systems - **Location:** Mountain View, CA, United States - **Experience:** Expert - **Salary:** $195,000.0 - $261,000.0 - **Contract:** Permanent contract - **Skills:** Systems Engineering, ARM Architecture, Computer Vision, C++ (Programming Language), Program Optimization, Nvidia CUDA, Computer Engineering, Linux, Memory Management, Firmware, Field-Programmable Gate Array (FPGA), Machine Learning, Real-Time Operating Systems, Tensorflow, Smart Devices, Management of Software Versions, Graphics Processing Unit (GPU), Pytorch, Information Technology, Low Latency, ONNX (Open Neural Network Exchange) Format, Machine Learning Operations, TensorRT, Data Pipelines - **Published:** August 4, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=e9ccdf00ec34f8d8 ## About the Role * 10+ years of professional software or systems engineering experience, including at least 2 years focused on deploying ML models to embedded or edge devices; Bachelor's or Master's degree in Computer Science, Electrical Engineering, Computer Engineering, or equivalent practical experience. * Very strong C++ proficiency; working knowledge of CUDA; hands-on experience with PyTorch and at least one edge inference runtime such as TensorFlow Lite, ONNX Runtime, or TensorRT. * Practical experience with model optimization techniques including post-training quantization, quantization-aware training, pruning, and distillation; demonstrated ability to profile and optimize for latency, memory, and power on constrained hardware. * Working knowledge of embedded or edge platforms such as NVIDIA Jetson, Qualcomm, ARM Cortex, or comparable NPUs and SoCs, and of Linux or an RTOS; solid grasp of computer architecture concepts relevant to inference including memory hierarchy, fixed-point arithmetic, and accelerator offload; domain experience in computer vision or sensor processing on device. You'll Stand Out * Hands-on experience deploying computer vision models for detection or tracking tasks on embedded or edge hardware. * Experience with NVIDIA Jetson specifically, including TensorRT optimization and deployment on Jetson platforms. * Background in defense, autonomous systems, or robotics where real-time reliability matters. * Experience building or contributing to model update and OTA delivery pipelines for deployed edge devices. ## Description We are looking for a Senior Embedded Machine Learning Engineer to own the end-to-end process of taking trained ML models and deploying them efficiently onto resource-constrained edge hardware. This role sits at the intersection of machine learning, embedded systems, and hardware engineering. You will integrate, convert, and optimize models to run within strict constraints on latency, memory, power, and thermal budget, and build the supporting C++ infrastructure that hosts them on device. You will partner closely with the CVML team who build the models, the embedded and firmware teams who own the device, and the product team who define performance targets. Success means models that are not just accurate in the lab but fast, small, and dependable in the field. What You'll Do * Apply quantization, pruning, knowledge distillation, operator fusion, and graph optimization to shrink models and reduce inference cost while protecting accuracy; convert trained models into edge-deployable formats using ONNX and TensorRT. * Profile inference on target accelerators including GPUs, NPUs, DSPs, and FPGAs; measure latency, throughput, memory footprint, and power consumption, then drive the changes needed to hit performance targets. * Design, write, and maintain the C++ application code that hosts inference on device, including pre- and post-processing pipelines, data and memory management, threading, and interfaces to the rest of the embedded system; ensure the combined model and C++ stack meets real-time constraints and fits within device memory budget. * Build test harnesses to verify on-device accuracy against reference results and catch regressions from optimization or quantization; contribute to tooling for packaging, versioning, and delivering model updates to deployed devices. * Set best practices for edge deployment, review designs and code, and mentor other engineers on optimization and embedded ML techniques; work closely with research, firmware, and product teams to set realistic performance targets and feed hardware constraints back into model design. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Intelligent Data Selection for Continual Learning of AI Functions](https://www.wearedevelopers.com/videos/367-intelligent-data-selection-for-continual-learning-of-ai-functions) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) - [From Perception to Autonomy: Building Agentic Edge AI Robots with ROS 2](https://www.wearedevelopers.com/videos/100295-from-perception-to-autonomy-building-agentic-edge-ai-robots-with-ros-2) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it)