> Markdown version of [/jobs/ext/3236971-embedded-ai-engineer](https://www.wearedevelopers.com/jobs/ext/3236971-embedded-ai-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Embedded AI Engineer - **Company:** Prophesee - **Location:** Paris, France - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Computing Platforms, Artificial Neural Networks, Computer Vision, C++ (Programming Language), Profiling, Code Review, Nvidia CUDA, Extract Transform Load (ETL), Memory Management, Linux on Embedded Systems, Python (Programming Language), Machine Learning, Software Architecture, Tensorflow, Software Engineering, Pytorch, Deep Learning, Git, TensorRT - **Published:** September 18, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=23cb73999ea9f5f7 ## About the Role * Embedded AI / model optimization - strong requirement - Hands-on experience optimizing machine learning or deep learning workloads for edge or constrained targets. * Deep learning frameworks - strong requirement - Strong practical knowledge of frameworks such as PyTorch, TensorFlow or equivalent. * C++ and Python - strong requirement - Ability to implement performance-sensitive inference code and supporting optimization / benchmarking tools. * Performance profiling - strong requirement - Ability to analyze latency, throughput, memory use and algorithmic complexity and identify bottlenecks. * Inference optimization - strong requirement - Experience with quantization, model conversion, graph optimization, kernel/runtime optimization or equivalent techniques. * Embedded Linux / ARM - required - Understanding of embedded Linux, ARM platforms, cross-compilation and constrained computing environments. * NVIDIA Jetson / TensorRT / CUDA - highly desirable - Hands-on deployment or optimization experience on NVIDIA edge AI platforms. * Computer vision / perception - desirable - Experience with detection, tracking, image processing, event-based vision or related perception pipelines. * Hardware-aware mindset - required - Ability to reason about CPU/GPU/NPU capabilities, memory hierarchy, data movement, power and target constraints. * Benchmarking & validation - required - Ability to build reproducible performance benchmarks while tracking model accuracy and functional regressions. * Software engineering practices - required - Git, testing, documentation, code review and maintainable production implementation. Language skills Fluency in French and English Soft skills Strong problem-solving skills, strong analytical skills. Flexible to dynamic environments and fast changing technologies. Passionate about technology. This person must work well with other engineers in a team environment. Good sense of autonomy. Must be pragmatic and self-motivated to complete a task even if it is outside of just the "well known" realm. "Can Do Attitude" is preferred. ## Description You will work where AI algorithms meet real computing constraints, helping PROPHESEE turn advanced perception into efficient edge implementations. The role offers the opportunity to optimize models and complete inference pipelines for demanding embedded targets where latency, memory, bandwidth and power consumption are as important as algorithmic accuracy. Mission Design, integrate and optimize AI / perception algorithms and models for deployment on constrained computing targets. The role sits between AI, software and embedded engineering and focuses on translating algorithmic performance into efficient production implementations that meet latency, compute, memory and power constraints. Role description * Design, adapt and optimize AI, deep learning and computer vision models for embedded and edge deployment. * Profile end-to-end inference pipelines to identify compute, memory, latency and bandwidth bottlenecks. * Convert or adapt models for target runtimes and accelerators while preserving required accuracy. * Implement and evaluate quantization, pruning, graph optimization, kernel optimization and other model-compression techniques where relevant. * Optimize pre-processing, post-processing and data movement around AI inference, not only the neural network itself. * Develop C++ and Python tooling for benchmarking, profiling, conversion, validation and deployment. * Contribute to software architecture and development, specifically around memory management, scheduling and real-time constraints. * Establish reproducible benchmarks and compare accuracy / latency / power / memory trade-offs across targets. * Support deployment on edge AI targets. Key deliverables * Optimized AI / perception models and inference pipelines for selected embedded targets. * Reproducible benchmark suite covering accuracy, latency, throughput, memory and, where available, power consumption. * Model-conversion and optimization flows for the selected inference runtimes and accelerators. * C++ and Python tools for profiling, validation, benchmarking and deployment. * Optimization reports documenting bottlenecks, trade-offs and improvements versus baseline implementations. * Deployment-ready integration with the software stack and associated technical documentation. Performance indicators * Achievement of agreed accuracy, latency, throughput, memory and power targets on selected hardware. * Performance improvement versus baseline model / pipeline implementations. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Localized Open Models in Production: What Builders Need to Know](https://www.wearedevelopers.com/videos/100270-localized-open-models-in-production-what-builders-need-to-know) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs)