Senior Firmware Engineer, Edge AI / NPU Runtime

TACIT, Inc.
San Francisco, CA, United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$150,000.0 - $200,000.0
Working hours
Regular working hours
Job source

Tech stack

Gesture Recognition Artificial Intelligence Extract Transform Load (ETL) Data Transformation Software Debugging Memory Management Embedded Software Firmware FreeRTOS Machine Learning Performance Tuning Real-Time Operating Systems
+10 more
Tensorflow Sensor Fusion Data Streaming Toolchain Concurrency Low Latency ONNX (Open Neural Network Exchange) Format Hardware Acceleration Machine Learning Operations Feature Extraction

Job description

We’re looking for a Senior Firmware Engineer, Edge AI / NPU Runtime to help architect, optimize, and ship next-generation neurotech hardware with production-grade on-device intelligence. You will own critical parts of the embedded AI stack, from realtime sensor acquisition through preprocessing, NPU/DSP-accelerated inference, postprocessing, telemetry, and product deployment.

This is a hands-on role for someone who wants to work close to the hardware while shaping the intelligence users experience in the product. You’ll help define how models run on-device, how sensor data moves through the system, and how we meet tight latency, reliability, and power budgets in real-world use.

What you’ll do

  • Edge AI & NPU Inference
  • Own deployment of ML models onto embedded targets using NPUs, DSPs, MCUs, or other hardware accelerators.
  • Integrate embedded inference runtimes, vendor NPU/DSP SDKs, and model deployment workflows into production firmware.
  • Optimize inference latency, memory footprint, throughput, power consumption, and accelerator utilization on production hardware.
  • Partner with ML teams on quantization, operator support, model architecture tradeoffs, calibration datasets, and accuracy/performance regressions.
  • Realtime Sensor-to-Inference Systems
  • Build realtime sensor-to-inference pipelines, including acquisition, timestamping, synchronization, preprocessing, feature extraction, model execution, and postprocessing.
  • Design low-latency data movement using DMA, interrupts, ring buffers, deterministic scheduling, and efficient memory layouts.
  • Support streaming inference patterns such as sliding windows, temporal models, event-driven execution, and continuous sensor processing.
  • Maintain inference quality and timing guarantees under real-world conditions such as sensor noise, clock drift, dropped samples, variable system load, and power-state transitions.
  • Power-Optimized Embedded Firmware
  • Optimize end-to-end energy per inference across sensing, preprocessing, model execution, postprocessing, and idle time.
  • Use low-power firmware techniques such as sleep states, duty cycling, subsystem power gating, clock scaling, batching/windowing, and dynamic power management.
  • Profile and improve power consumption across sensors, CPU, NPU/DSP, memory, and supporting firmware infrastructure.
  • Product Quality & Debugging
  • Bring up and debug firmware across sensors, accelerators, power systems, embedded compute, and production hardware.
  • Use lab tools, traces, logs, telemetry, and instrumentation to root-cause complex embedded system issues.
  • Translate product and customer experience goals into concrete latency, reliability, responsiveness, and power targets.
  • Build diagnostics, validation hooks, and performance benchmarks to ensure reliable real-world edge inference behavior.

Requirements

Do you have experience in System performance optimization?, * 5+ years of experience in embedded firmware, embedded systems, or edge ML systems.

  • Strong C/C++/Rust experience on resource-constrained embedded platforms.
  • Experience with RTOS-based systems such as FreeRTOS, Zephyr, ThreadX, or similar.
  • Experience deploying or optimizing ML inference on embedded targets, NPUs, DSPs, MCUs, or edge SoCs.
  • Strong understanding of realtime embedded systems, including DMA, interrupts, concurrency, memory management, and low-latency data movement.
  • Experience optimizing embedded systems for latency, memory footprint, throughput, and power consumption.
  • Hands-on debugging and bring-up experience across embedded hardware and firmware systems, with strong cross-functional communication across firmware, ML, electrical, software, and product teams., * Experience with embedded inference runtimes, deployment toolchains, or edge AI SoCs/accelerators such as TensorFlow Lite Micro, ONNX Runtime, CMSIS-NN, Qualcomm QNN/SNPE, ARM Ethos-U/Vela, TVM, ExecuTorch, Qualcomm, ARM, Cadence/Tensilica, Syntiant, Ambiq, Nordic, NXP, ST, Hailo, Google Edge TPU, or similar.
  • Experience with quantized inference, fixed-point math, SIMD/DSP optimization, accelerator programming, or model conversion workflows.
  • Experience with streaming or time-series ML workloads such as biosignals, sensor fusion, audio, gesture recognition, keyword spotting, or other realtime inference systems.
  • Experience shipping battery-powered consumer electronics, wearable, neurotech, AR/VR, robotics, camera, IoT, or other embedded AI products.

Benefits & conditions

2.32.3 out of 5 stars San Francisco, CA $150,000 - $200,000 a year - Full-time, Pulled from the full job description

  • Health insurance
  • 401(k) matching
  • Vision insurance
  • 401(k) 3% Match
  • Dental insurance
  • Unlimited paid time off
  • Visa sponsorship, * Competitive equity package
  • Comprehensive medical, dental, and vision insurance
  • Company size: 20-30 people
  • Unlimited PTO
  • Visa sponsorship
  • 3% 401k matching

Compensation Range: $150K - $200K

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:56 min

Optimizing AI processing capabilities for edge microcontrollers

Stephan Gillich Stephan Gillich +3 · WWC 2024

2:19 min

Orchestrating over-the-air firmware updates for vehicle modules

Denis Grahovac · WWC 2021

3:12 min

Integrating toolchains to improve engineering efficiency and culture

Amir Friedman Amir Friedman +2 · WWC 2025

3:40 min

Hardware constraints and core challenges in edge environments

Ekaterina Sirazitdinova · WWC 2023

2:20 min

Utilizing custom firmware for variable torque manipulation

Daniel Meilak Daniel Meilak +1 · WWC Europe 2026

2:11 min

Promoting an open toolchain to maximize developer choice

Anu Bharadwaj Anu Bharadwaj · WWC 2023

Videos

See all

Related articles

See all