> Markdown version of [/jobs/ext/1105458-system-software-engineer-gpu-accelerated-compute](https://www.wearedevelopers.com/jobs/ext/1105458-system-software-engineer-gpu-accelerated-compute). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # System Software Engineer - GPU & Accelerated Compute - **Company:** Sunday Inc - **Location:** Redwood City, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), C++ (Programming Language), Profiling, Nvidia CUDA, Data Transmissions, Extract Transform Load (ETL), Linux, Memory Management, Performance Tuning, Software Engineering, System Software, Graphics Processing Unit (GPU), Gpu Programming, Enterprise Integration, Decoding - **Published:** June 9, 2026 - **Apply:** https://www.dice.com/job-detail/b8f2ea0f-2211-457b-ae23-24fa1a30e3dc ## About the Role * 2+ years of experience developing gpu systems software * Strong proficiency in CUDA and a systems language such as C++, C, or Rust * Solid understanding of GPU architecture, GPU workloads, and the tradeoffs involved in time-slicing and sharing the device across users * Hands-on experience with the CUDA ecosystem: CUDA runtime API, CUDA Graphs, and CUDA IPC * Familiarity with GPU sharing mechanisms such as MPS and MIG * Experience with GPU profiling tools such as Nsight Systems and Nsight Compute * Solid Linux fundamentals: scheduling, IPC, memory management, and performance tuning Nice to Have * Contributions to CUDA libraries or other GPU programming libraries * Experience with camera pipeline integration and NVDEC/NVENC * Experience optimizing model inference on embedded GPU platforms (e.g., Jetson) * Experience with observability and tracing for GPU-accelerated workloads ## Description At Sunday, we're developing personal robots to reclaim the hours lost to repetitive tasks. We're focused on an ambitious goal to make generalized robots broadly accessible, enabling households to take back quality time. We have spent the last 18 months building a talented team, securing capital, and validating our technology. We are now seeking passionate individuals to join us in the next phase of our growth. If you are ready to apply your skills to the forefront of robotics innovation, we'd love to hear from you. What to Expect The ML & Robotics Infra team builds the foundational systems that every part of our robot perception, ML, controls and behavior runs on, and the developer infrastructure that lets us build, ship, and update that software quickly and safely on every robot in the fleet. As a System Software Engineer on ML & Robotics Infra focused on GPU and accelerated compute, you'll own how every accelerated workload on the robot from model inference, SLAM/perception, and more gets data, gets scheduled and runs efficiently on shared compute. You'll work alongside teammates who own the runtime and our build and delivery infrastructure, and you'll partner cross-functionally with ML, SLAM/Perception, Controls and Hardware teams to ensure the GPU is a first-class, well-utilized resource that meets the latency and throughput requirements of a real-time robotic system operating in the home. What You'll Do You'll own and contribute to the accelerated compute layer of the ML & Robotics Infra, including: * Efficient model execution and switching: Reduce gpu kernel launch overheads and make swapping between models on the same device fast and predictable * GPU scheduling and time-slicing: Arbitrate GPU access across concurrent users (model inference, SLAM, and other robotics applications) with predictable latency * Camera pipeline: Drive low-latency transfer of camera frames into GPU memory, integrating with HW accelerate encode/decode (NVDEC/NVENC) where appropriate * CPU GPU data transfer: Build efficient, low-overhead data movement between host and device, including pinned memory, zero-copy paths, and asynchronous transfer patterns * CPU/GPU synchronization: Design synchronization primitives and patterns that minimize stalls and keep inference pipelines full ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Profiling Symfony & PHP apps with Blackfire](https://www.wearedevelopers.com/videos/265-profiling-symfony-php-apps-with-blackfire) - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Challenges and Solutions for Efficient, Large-Scale Video Analysis](https://www.wearedevelopers.com/videos/2022-challenges-and-solutions-for-efficient-large-scale-video-analysis) - [Accelerating Python on GPUs](https://www.wearedevelopers.com/videos/1521-accelerating-python-on-gpus) ## Related Articles - [What’s the latest in NVIDIA CUDA Python](https://www.wearedevelopers.com/magazine/568-what-s-the-latest-in-nvidia-cuda-python) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How software is steering vehicle technology](https://www.wearedevelopers.com/magazine/515-how-software-is-steering-vehicle-technology) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Why Attend a Developer Event in 2026?](https://www.wearedevelopers.com/magazine/688-why-attend-a-developer-event-in-2026) - [Top 6 Hackathons for Developers in 2023](https://www.wearedevelopers.com/magazine/263-top-6-hackathons-for-developers-in-2023)