Staff Software Engineer, Deep Learning Acceleration
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+6 more
Job description
Experteer Overview In this role you drive the performance optimization of deep learning networks used in Aurora’s autonomous vehicle systems. You will analyze and optimize software architecture, latency, and deployment both on-vehicle and in data centers. You’ll collaborate with cross-functional teams to scale DL workloads, improving efficiency and reliability of self-driving technology. Your work helps make transportation safer and more accessible at scale. This is a fast-paced, impact-oriented role in a company solving complex, meaningful problems in mobility. Compensation / Benefits * Conduct performance analysis and optimization of Deep Learning networks running on the AV * Improve software architecture, system performance, and latency for deep learning applications * Deploy DL models on the AV and for large-scale data center training * Troubleshoot performance issues using profiling and roofline model techniques * Collaborate with cross-functional teams to enhance self-driving efficiency Tasks * 5+ years of software engineering experience * BS, MS, or PhD in Computer Science or related field * Proficiency in CUDA, C++, and Python * Experience in high-performance computing and parallel programming * Skill with performance analysis tools (NVIDIA Nsight Systems, Nsight Compute) and roofline model * Hands-on DL/ML framework experience (PyTorch, TensorFlow) for model deployment * Understanding of CV and transformer-based architectures * Strong analytical and troubleshooting skills * Ability to learn new technologies quickly in a fast-paced environment * Experience working with large codebases and cross-functional teams * Linux/Unix proficiency Key requirements * annual bonus * equity compensation * benefits * hybrid work environment * in-office 3 days per week
Requirements
codebases Tasks * 5+ years of software engineering experience * BS, MS, or PhD in Computer Science or related field * Proficiency in CUDA, C++, and Python * Experience in high-performance computing and parallel programming * Skill with performance analysis tools (NVIDIA Nsight Systems, Nsight Compute) and roofline model * Hands-on DL/ML framework experience (PyTorch, TensorFlow) for model deployment * Understanding of CV and transformer-based architectures * Strong analytical and troubleshooting skills * Ability to learn new technologies quickly in a fast-paced environment * Experience working with large codebases and cross-functional teams * Linux/Unix proficiency Key requirements * annual bonus * equity compensation * benefits * hybrid work environment * in-office 3 days per week
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on us.experteer.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
MLOps And AI Driven Development
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence
How to Become an AI Engineer
Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production