> Markdown version of [/jobs/ext/177758-edge-ml-embedded-engineer](https://www.wearedevelopers.com/jobs/ext/177758-edge-ml-embedded-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Edge ML / Embedded Engineer - **Company:** Data Inc - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, C++ (Programming Language), Cloud Database, Computer Engineering, Concurrent Computing, Data Retrieval, Data Stores, Memory Management, Linux on Embedded Systems, Hardware Interface Design, Python (Programming Language), Machine Learning, Tensorflow, Message Oriented Middleware, Data Streaming, Information Technology, ONNX (Open Neural Network Exchange) Format, Real Time Data, Operational Systems, Speech Synthesis, Multiaccess Edge Computing, Industrial Software - **Published:** May 31, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=db11220dd22d55a5 ## About the Role Do you have experience in Technical report writing?, * Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent professional experience in embedded systems or edge computing. * 5+ years of hands-on experience in embedded systems engineering, edge computing, or on-device machine learning, with demonstrated work on constrained hardware environments. * Expert-level proficiency with at least one edge ML inference framework: TensorFlow Lite, ONNX Runtime, llama.cpp, or equivalent. Experience optimizing and quantizing models for CPU-only inference is required. * Strong understanding of memory management, real-time data stream handling, and concurrent processing in resource-constrained environments. Experience with C++, Rust, or Python with tight memory management is strongly preferred. * Experience with embedded Linux or equivalent OS environments, including ARM-based processors, limited RAM, and environments without GPU availability. * Familiarity with real-time data ingestion from hardware interfaces or industrial systems - including serial protocols, message bus architectures, or event-driven pipelines at the edge. * AWS familiarity preferred, specifically IoT Greengrass as a candidate edge runtime and IoT Core for device-to-cloud connectivity. Hands-on implementation experience is not required but direct familiarity strengthens the candidate's ability to evaluate candidate architectures. * Experience with voice-to-text or text-to-speech pipelines in offline or low-connectivity environments is a plus. * Comfortable operating in a Phase 0 discovery and feasibility mode - producing assessment findings, ADRs, and a constrained demonstrator rather than production-ready software. * Strong written communication skills with the ability to document hardware constraint findings, framework evaluations, and architectural trade-offs in formats usable by both technical architects and client stakeholders. * Experience working in consulting or client-facing project environments is preferred. If you are an embedded systems or edge ML engineer who is energized by early-stage technical discovery work - evaluating what is feasible before committing to what will be built - and you bring deep hands-on experience making AI work on hardware that was never designed for it, we invite you to apply. ## Description * Assess target edge hardware against the requirements of an on-device inference loop: evaluate processor architecture, available memory, OS and runtime environment, and whether candidate edge runtimes (such as IoT Greengrass or equivalent) can be supported. * Evaluate candidate edge inference frameworks for CPU-only SLM deployment - including TensorFlow Lite, ONNX Runtime, llama.cpp, and equivalents - assessing quantization approaches, inference latency, and memory footprint against feasibility targets confirmed during discovery. * Assess real-time data ingestion feasibility from operational subsystem interfaces, evaluating candidate patterns for consuming concurrent data streams within the memory and compute constraints of the target hardware. * Design and evaluate local data store options for the on-device SLM context, including storage formats, retrieval latency, and update mechanisms appropriate for the edge environment. * Build a constrained feasibility demonstrator on laptop or workstation hardware using simulated data feeds. The demonstrator validates the interaction model and core architectural approach - it is not a production prototype and does not connect to operational systems. * Implement a small number of scoped interaction flows in the demonstrator, integrating the voice interface pipeline with the SLM inference and local data retrieval components as agreed through the engagement scope. * Collaborate with the AI/ML Architect on SLM selection, domain restriction approach, and inference pipeline design - providing hardware and runtime constraint inputs that shape what is architecturally feasible. * Collaborate with the AWS Solutions Architect on the edge-to-cloud data channel, identifying what can realistically be buffered and transmitted from a constrained edge device under variable connectivity conditions. * Document hardware assessment findings, framework evaluations, and architectural trade-offs as Architecture Decision Records (ADRs) with explicit rationale. Clearly flag where recommendations are conditional on hardware or interface specifications not yet confirmed. * Communicate technical constraints and feasibility findings clearly to both technical architects and non-technical client stakeholders throughout the engagement. ## Related Videos - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) - [Machine learning in the browser with TensorFlowjs](https://www.wearedevelopers.com/videos/155-machine-learning-in-the-browser-with-tensorflowjs) - [Fault Tolerance and Consistency at Scale: Harnessing the Power of Distributed SQL Databases](https://www.wearedevelopers.com/videos/1520-fault-tolerance-and-consistency-at-scale-harnessing-the-power-of-distributed-sql-databases) - [Tomorrow's cloud data platforms - fully managed database-as-a-service (DBaaS)](https://www.wearedevelopers.com/videos/254-tomorrow-s-cloud-data-platforms-fully-managed-database-as-a-service-dbaas) - [From Perception to Autonomy: Building Agentic Edge AI Robots with ROS 2](https://www.wearedevelopers.com/videos/100295-from-perception-to-autonomy-building-agentic-edge-ai-robots-with-ros-2) - [Overview of Machine Learning in Python](https://www.wearedevelopers.com/videos/840-overview-of-machine-learning-in-python) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Dev Digest 132 - Binging WADFlix?](https://www.wearedevelopers.com/magazine/473-dev-digest-132-binging-wadflix)