Software Engineer, Edge Systems & Runtime (C++), Materra
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+9 more
Job description
to join our team, focusing on high-performance C++ software and on-device machine learning execution. In this role, you will work on low-latency software pipelines that process multi-sensor data streams and execute deep learning models on edge GPUs.
Your primary focus will be optimizing our production C++ runtime infrastructure, refining memory usage, improving data streaming between host and device, and tuning execution engines to squeeze out low-latency optimizations within single-digit milliseconds., * Profile and optimize established C++ software pipelines and GPU runtime engines to reduce latency and maximize hardware utilization.
- Refine host-to-device streaming pipelines using efficient memory management patterns (e.g., pinned memory, zero-copy buffers, lock-free queues).
- Deploy, update, and optimize deep learning model execution runtimes (e.g., TensorRT, CUDA, ONNX Runtime) on edge GPUs.
- Enhance asynchronous C++ data streams routing multi-sensor inputs into model inference loops.
- Collaborate with machine learning engineers to ensure smooth handoffs from training pipelines to production C++ runtimes.
Requirements
- Education: Degree in Computer Science, Computer Engineering, Electrical Engineering, Robotics, or a related technical field.
- Modern C++ Proficiency: 3+ years expertise in modern C++ (C++17/20), including multithreading, asynchronous I/O, IPC, and manual/smart memory management.
- On-Device ML Deployment: 3+ years hands-on experience profiling, compiling, and optimizing deep learning models for edge hardware using TensorRT, CUDA, ONNX, or similar acceleration frameworks.
- Systems Optimization: Proven experience navigating, profiling, and optimizing C++ execution paths under strict, low-latency performance constraints.
Preferred Skills
- Sensor & Hardware Integration: Familiarity with hardware communication APIs, SDKs, or sensor streams (e.g., GigE Vision, GenICam, frame-grabbers, or point clouds).
- Domain Experience: 3+ years exposure to real-time perception systems in robotics, industrial automation, autonomous vehicles, or video processing engines.
- Scripting & Automation: Proficiency in Python for rapid prototyping and system integration.
- Advanced Acceleration: Experience writing custom CUDA kernels, custom TensorRT plugins, or working with real-time OS extensions.
Benefits & conditions
The US base salary range for this full-time position is $166,000 - $244,000 + bonus + equity + benefits. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training. Your recruiter can share more about the specific salary range for your location during the hiring process.
Please note that the compensation details listed in US role postings reflect the base salary only, and do not include bonus, equity, or benefits.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on dejobs.orgGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Dev Digest 120 - Apple and peers
Dev Digest 129 - Now that's what I call private data!
Dev Digest 102 - Race conditions
MLOps And AI Driven Development