> Markdown version of [/jobs/ext/1946192-staff-software-engineer-deep-learning-acceleration](https://www.wearedevelopers.com/jobs/ext/1946192-staff-software-engineer-deep-learning-acceleration). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Software Engineer, Deep Learning Acceleration - **Company:** Aurora Innovation, Inc. - **Location:** Mountain View, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Big Data, C++ (Programming Language), Nvidia CUDA, Data Centers, Linux, Python (Programming Language), Performance Tuning, Software Architecture, Tensorflow, Software Engineering, High Performance Computing, Pytorch, Deep Learning, Parallel Computation, Information Technology, Codebase, Machine Learning Operations, GPT - **Published:** August 6, 2026 - **Apply:** https://us.experteer.com/career/view-jobs/staff-software-engineer-deep-learning-acceleration-mountain-view-ca-usa-58818664 ## About the Role codebases Tasks * 5+ years of software engineering experience * BS, MS, or PhD in Computer Science or related field * Proficiency in CUDA, C++, and Python * Experience in high-performance computing and parallel programming * Skill with performance analysis tools (NVIDIA Nsight Systems, Nsight Compute) and roofline model * Hands-on DL/ML framework experience (PyTorch, TensorFlow) for model deployment * Understanding of CV and transformer-based architectures * Strong analytical and troubleshooting skills * Ability to learn new technologies quickly in a fast-paced environment * Experience working with large codebases and cross-functional teams * Linux/Unix proficiency Key requirements * annual bonus * equity compensation * benefits * hybrid work environment * in-office 3 days per week ## Description Experteer Overview In this role you drive the performance optimization of deep learning networks used in Aurora's autonomous vehicle systems. You will analyze and optimize software architecture, latency, and deployment both on-vehicle and in data centers. You'll collaborate with cross-functional teams to scale DL workloads, improving efficiency and reliability of self-driving technology. Your work helps make transportation safer and more accessible at scale. This is a fast-paced, impact-oriented role in a company solving complex, meaningful problems in mobility. Compensation / Benefits * Conduct performance analysis and optimization of Deep Learning networks running on the AV * Improve software architecture, system performance, and latency for deep learning applications * Deploy DL models on the AV and for large-scale data center training * Troubleshoot performance issues using profiling and roofline model techniques * Collaborate with cross-functional teams to enhance self-driving efficiency Tasks * 5+ years of software engineering experience * BS, MS, or PhD in Computer Science or related field * Proficiency in CUDA, C++, and Python * Experience in high-performance computing and parallel programming * Skill with performance analysis tools (NVIDIA Nsight Systems, Nsight Compute) and roofline model * Hands-on DL/ML framework experience (PyTorch, TensorFlow) for model deployment * Understanding of CV and transformer-based architectures * Strong analytical and troubleshooting skills * Ability to learn new technologies quickly in a fast-paced environment * Experience working with large codebases and cross-functional teams * Linux/Unix proficiency Key requirements * annual bonus * equity compensation * benefits * hybrid work environment * in-office 3 days per week ## Related Videos - [WWC24 - Ankit Patel - Unlocking the Future Breakthrough Application Performance and Capabilities with NVIDIA](https://www.wearedevelopers.com/videos/920-wwc24-ankit-patel-unlocking-the-future-breakthrough-application-performance-and-capabilities-with-nvidia) - [ Evaluating AI models for code comprehension](https://www.wearedevelopers.com/videos/1462-evaluating-ai-models-for-code-comprehension) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)