Software Engineer, Edge Systems & Runtime (C++), Materra

Google LLC
Mountain View, CA, United States
12 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Compensation
$166,000.0 - $244,000.0
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) C++ (Programming Language) Profiling Nvidia CUDA Computer Engineering Memory Management Hardware Design Python (Programming Language) Machine Learning Rapid Prototyping Process Real-Time Operating Systems Data Streaming
+9 more
Systems Integration Video Editing Graphics Processing Unit (GPU) Deep Learning Information Technology ONNX (Open Neural Network Exchange) Format TensorRT Multiaccess Edge Computing C++14

Job description

to join our team, focusing on high-performance C++ software and on-device machine learning execution. In this role, you will work on low-latency software pipelines that process multi-sensor data streams and execute deep learning models on edge GPUs.

Your primary focus will be optimizing our production C++ runtime infrastructure, refining memory usage, improving data streaming between host and device, and tuning execution engines to squeeze out low-latency optimizations within single-digit milliseconds., * Profile and optimize established C++ software pipelines and GPU runtime engines to reduce latency and maximize hardware utilization.

  • Refine host-to-device streaming pipelines using efficient memory management patterns (e.g., pinned memory, zero-copy buffers, lock-free queues).
  • Deploy, update, and optimize deep learning model execution runtimes (e.g., TensorRT, CUDA, ONNX Runtime) on edge GPUs.
  • Enhance asynchronous C++ data streams routing multi-sensor inputs into model inference loops.
  • Collaborate with machine learning engineers to ensure smooth handoffs from training pipelines to production C++ runtimes.

Requirements

  • Education: Degree in Computer Science, Computer Engineering, Electrical Engineering, Robotics, or a related technical field.
  • Modern C++ Proficiency: 3+ years expertise in modern C++ (C++17/20), including multithreading, asynchronous I/O, IPC, and manual/smart memory management.
  • On-Device ML Deployment: 3+ years hands-on experience profiling, compiling, and optimizing deep learning models for edge hardware using TensorRT, CUDA, ONNX, or similar acceleration frameworks.
  • Systems Optimization: Proven experience navigating, profiling, and optimizing C++ execution paths under strict, low-latency performance constraints.

Preferred Skills

  • Sensor & Hardware Integration: Familiarity with hardware communication APIs, SDKs, or sensor streams (e.g., GigE Vision, GenICam, frame-grabbers, or point clouds).
  • Domain Experience: 3+ years exposure to real-time perception systems in robotics, industrial automation, autonomous vehicles, or video processing engines.
  • Scripting & Automation: Proficiency in Python for rapid prototyping and system integration.
  • Advanced Acceleration: Experience writing custom CUDA kernels, custom TensorRT plugins, or working with real-time OS extensions.

Benefits & conditions

The US base salary range for this full-time position is $166,000 - $244,000 + bonus + equity + benefits. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training. Your recruiter can share more about the specific salary range for your location during the hiring process.

Please note that the compensation details listed in US role postings reflect the base salary only, and do not include bonus, equity, or benefits.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · WWC Europe 2026

4:52 min

Essential phases in building and refining language models

Anshul Jindal Anshul Jindal +1 · WWC 2025

47 sec

Profiling native execution calls with async-profiler

Gonzalo Ortiz Jaureguizar Gonzalo Ortiz Jaureguizar · WWC Europe 2026

2:32 min

Core libraries driving inference engines and multi-GPU networking

Adolf Hohl Adolf Hohl · WWC 2024

1:50 min

Overview of the Edge AI ecosystem and tech stack

Maxim Salnikov Maxim Salnikov · WWC 2025

2:17 min

Comparing code profiling with surface level monitoring

Jérôme Vieilledent · LIVE

Videos

See all

Related articles

See all