> Markdown version of [/jobs/ext/1501901-software-engineer-edge-systems-runtime-c-materra](https://www.wearedevelopers.com/jobs/ext/1501901-software-engineer-edge-systems-runtime-c-materra). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, Edge Systems & Runtime (C++), Materra - **Company:** Google LLC - **Location:** Mountain View, CA, United States - **Experience:** Experienced - **Salary:** $166,000.0 - $244,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), C++ (Programming Language), Profiling, Nvidia CUDA, Computer Engineering, Memory Management, Hardware Design, Python (Programming Language), Machine Learning, Rapid Prototyping Process, Real-Time Operating Systems, Data Streaming, Systems Integration, Video Editing, Graphics Processing Unit (GPU), Deep Learning, Information Technology, ONNX (Open Neural Network Exchange) Format, TensorRT, Multiaccess Edge Computing, C++14 - **Published:** July 30, 2026 - **Apply:** https://dejobs.org/x/x/EC3063C606474F439CA3E5C278C5BA49/job/ ## About the Role * Education: Degree in Computer Science, Computer Engineering, Electrical Engineering, Robotics, or a related technical field. * Modern C++ Proficiency: 3+ years expertise in modern C++ (C++17/20), including multithreading, asynchronous I/O, IPC, and manual/smart memory management. * On-Device ML Deployment: 3+ years hands-on experience profiling, compiling, and optimizing deep learning models for edge hardware using TensorRT, CUDA, ONNX, or similar acceleration frameworks. * Systems Optimization: Proven experience navigating, profiling, and optimizing C++ execution paths under strict, low-latency performance constraints. Preferred Skills * Sensor & Hardware Integration: Familiarity with hardware communication APIs, SDKs, or sensor streams (e.g., GigE Vision, GenICam, frame-grabbers, or point clouds). * Domain Experience: 3+ years exposure to real-time perception systems in robotics, industrial automation, autonomous vehicles, or video processing engines. * Scripting & Automation: Proficiency in Python for rapid prototyping and system integration. * Advanced Acceleration: Experience writing custom CUDA kernels, custom TensorRT plugins, or working with real-time OS extensions. ## Description to join our team, focusing on high-performance C++ software and on-device machine learning execution. In this role, you will work on low-latency software pipelines that process multi-sensor data streams and execute deep learning models on edge GPUs. Your primary focus will be optimizing our production C++ runtime infrastructure, refining memory usage, improving data streaming between host and device, and tuning execution engines to squeeze out low-latency optimizations within single-digit milliseconds., * Profile and optimize established C++ software pipelines and GPU runtime engines to reduce latency and maximize hardware utilization. * Refine host-to-device streaming pipelines using efficient memory management patterns (e.g., pinned memory, zero-copy buffers, lock-free queues). * Deploy, update, and optimize deep learning model execution runtimes (e.g., TensorRT, CUDA, ONNX Runtime) on edge GPUs. * Enhance asynchronous C++ data streams routing multi-sensor inputs into model inference loops. * Collaborate with machine learning engineers to ensure smooth handoffs from training pipelines to production C++ runtimes. ## Related Videos - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Profiling Symfony & PHP apps with Blackfire](https://www.wearedevelopers.com/videos/265-profiling-symfony-php-apps-with-blackfire) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Enhancing Workload Security in Kubernetes](https://www.wearedevelopers.com/videos/356-enhancing-workload-security-in-kubernetes) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 129 - Now that's what I call private data!](https://www.wearedevelopers.com/magazine/468-dev-digest-129-now-that-s-what-i-call-private-data) - [Dev Digest 102 - Race conditions](https://www.wearedevelopers.com/magazine/386-dev-digest-102-race-conditions) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this)