> Markdown version of [/jobs/ext/879194-senior-firmware-engineer-edge-ai-npu-runtime](https://www.wearedevelopers.com/jobs/ext/879194-senior-firmware-engineer-edge-ai-npu-runtime). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Firmware Engineer, Edge AI / NPU Runtime - **Company:** TACIT, Inc. - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Salary:** $150,000.0 - $200,000.0 - **Contract:** Permanent contract - **Skills:** Gesture Recognition, Artificial Intelligence, Extract Transform Load (ETL), Data Transformation, Software Debugging, Memory Management, Embedded Software, Firmware, FreeRTOS, Machine Learning, Performance Tuning, Real-Time Operating Systems, Tensorflow, Sensor Fusion, Data Streaming, Toolchain, Concurrency, Low Latency, ONNX (Open Neural Network Exchange) Format, Hardware Acceleration, Machine Learning Operations, Feature Extraction - **Published:** June 4, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=8d6c923910cfdf94 ## About the Role Do you have experience in System performance optimization?, * 5+ years of experience in embedded firmware, embedded systems, or edge ML systems. * Strong C/C++/Rust experience on resource-constrained embedded platforms. * Experience with RTOS-based systems such as FreeRTOS, Zephyr, ThreadX, or similar. * Experience deploying or optimizing ML inference on embedded targets, NPUs, DSPs, MCUs, or edge SoCs. * Strong understanding of realtime embedded systems, including DMA, interrupts, concurrency, memory management, and low-latency data movement. * Experience optimizing embedded systems for latency, memory footprint, throughput, and power consumption. * Hands-on debugging and bring-up experience across embedded hardware and firmware systems, with strong cross-functional communication across firmware, ML, electrical, software, and product teams., * Experience with embedded inference runtimes, deployment toolchains, or edge AI SoCs/accelerators such as TensorFlow Lite Micro, ONNX Runtime, CMSIS-NN, Qualcomm QNN/SNPE, ARM Ethos-U/Vela, TVM, ExecuTorch, Qualcomm, ARM, Cadence/Tensilica, Syntiant, Ambiq, Nordic, NXP, ST, Hailo, Google Edge TPU, or similar. * Experience with quantized inference, fixed-point math, SIMD/DSP optimization, accelerator programming, or model conversion workflows. * Experience with streaming or time-series ML workloads such as biosignals, sensor fusion, audio, gesture recognition, keyword spotting, or other realtime inference systems. * Experience shipping battery-powered consumer electronics, wearable, neurotech, AR/VR, robotics, camera, IoT, or other embedded AI products. ## Description We're looking for a Senior Firmware Engineer, Edge AI / NPU Runtime to help architect, optimize, and ship next-generation neurotech hardware with production-grade on-device intelligence. You will own critical parts of the embedded AI stack, from realtime sensor acquisition through preprocessing, NPU/DSP-accelerated inference, postprocessing, telemetry, and product deployment. This is a hands-on role for someone who wants to work close to the hardware while shaping the intelligence users experience in the product. You'll help define how models run on-device, how sensor data moves through the system, and how we meet tight latency, reliability, and power budgets in real-world use. What you'll do * Edge AI & NPU Inference + Own deployment of ML models onto embedded targets using NPUs, DSPs, MCUs, or other hardware accelerators. + Integrate embedded inference runtimes, vendor NPU/DSP SDKs, and model deployment workflows into production firmware. + Optimize inference latency, memory footprint, throughput, power consumption, and accelerator utilization on production hardware. + Partner with ML teams on quantization, operator support, model architecture tradeoffs, calibration datasets, and accuracy/performance regressions. * Realtime Sensor-to-Inference Systems + Build realtime sensor-to-inference pipelines, including acquisition, timestamping, synchronization, preprocessing, feature extraction, model execution, and postprocessing. + Design low-latency data movement using DMA, interrupts, ring buffers, deterministic scheduling, and efficient memory layouts. + Support streaming inference patterns such as sliding windows, temporal models, event-driven execution, and continuous sensor processing. + Maintain inference quality and timing guarantees under real-world conditions such as sensor noise, clock drift, dropped samples, variable system load, and power-state transitions. * Power-Optimized Embedded Firmware + Optimize end-to-end energy per inference across sensing, preprocessing, model execution, postprocessing, and idle time. + Use low-power firmware techniques such as sleep states, duty cycling, subsystem power gating, clock scaling, batching/windowing, and dynamic power management. + Profile and improve power consumption across sensors, CPU, NPU/DSP, memory, and supporting firmware infrastructure. * Product Quality & Debugging + Bring up and debug firmware across sensors, accelerators, power systems, embedded compute, and production hardware. + Use lab tools, traces, logs, telemetry, and instrumentation to root-cause complex embedded system issues. + Translate product and customer experience goals into concrete latency, reliability, responsiveness, and power targets. + Build diagnostics, validation hooks, and performance benchmarks to ensure reliable real-world edge inference behavior. ## Related Videos - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) - [Speeding up Web Apps performance with WebAssembly and Emscripten](https://www.wearedevelopers.com/videos/1985-speeding-up-web-apps-performance-with-webassembly-and-emscripten) - [Playing Pong on a shoulder press machine](https://www.wearedevelopers.com/videos/100140-playing-pong-on-a-shoulder-press-machine) - [From Perception to Autonomy: Building Agentic Edge AI Robots with ROS 2](https://www.wearedevelopers.com/videos/100295-from-perception-to-autonomy-building-agentic-edge-ai-robots-with-ros-2) - [The Power of Developer Communities](https://www.wearedevelopers.com/videos/1109-the-power-of-developer-communities) - [Agent Smith Gets Hardware: Autonomous IoT Hacking From Debug Port to Cloud API](https://www.wearedevelopers.com/videos/100258-agent-smith-gets-hardware-autonomous-iot-hacking-from-debug-port-to-cloud-api) ## Related Articles - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Dev Digest 138 - Are you secure about this?](https://www.wearedevelopers.com/magazine/486-dev-digest-138-are-you-secure-about-this) - [Dev Digest 118 - not a total recall](https://www.wearedevelopers.com/magazine/452-dev-digest-118-not-a-total-recall) - [Dev Digest 132 - Binging WADFlix?](https://www.wearedevelopers.com/magazine/473-dev-digest-132-binging-wadflix)