Staff/Sr. Staff Software Engineer, AI Software Tools (Onsite)

Qualcomm
Raleigh, NC, United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
2 years minimum
Compensation
$158,400.0 - $237,600.0
Working hours
Regular working hours
Job source

Tech stack

Adobe InDesign Agile Methodology Artificial Intelligence Algorithm Design Systems Engineering CMake Program Optimization Software Quality Code Review Information Systems Computer Engineering Software Debugging
+25 more
Memory Management Embedded Software Graph Theory Python (Programming Language) Logical Volume Manager Machine Learning Performance Tuning Real-Time Operating Systems Software Tools Software Engineering Software Systems Pytorch Delivery Pipeline Large Language Models Deep Learning Reliability of Systems Generative AI Git Build Management Information Technology Low Latency Optimization Algorithms ONNX (Open Neural Network Exchange) Format HuggingFace Decoding

Job description

As a Staff/Sr. Staff Software Engineer in the Qualcomm AI Stack SDK Software team, you will design, develop, and deliver advanced AI/ML software solutions for Generative AI inference on Snapdragon platforms. This role focuses on model optimization, quantization, graph transformations, and runtime execution for modern AI architectures including LLMs, LVMs, and LMMs., You will work at the intersection of machine learning algorithms, inference optimization, graph lowering, and systems software, contributing directly to the Qualcomm AI Stack SDK (QAIRT), and associated tools, including delegates support for ONNX Runtime, Executorch and TFLite/LiteRT frameworks. You will collaborate with amazing engineers from different teams across multiple locations like ML Research, AI accelerator HW/SW teams, Product Management, Program Management, and QA to drive features from concept to production., * Convert, optimize, and deploy AI models from PyTorch and ONNX frameworks for efficient inference on Snapdragon platforms.

  • Design and implement graph transformations, graph lowering, and optimization techniques within AI runtime environments such as ONNX Runtime, ExecuTorch and Qualcomm AI Stack SDK.

  • Apply knowledge of quantization and performance optimization to improve latency, throughput, memory usage, and power efficiency.

  • Work at the forefront of Generative AI, understanding advanced algorithms such as attention mechanisms, Mixture-of-Experts (MoE), Low Rank Adapter (LoRA) and emerging inference optimization techniques (e.g., Speculative Decoding etc.).

  • Collaborate with ML Research teams to prototype and productize new features and techniques into SDK solutions.

  • Debug complex issues across models, runtime, OS, compiler, and hardware layers, working closely with QA and customer teams.

  • Design, implement, and deliver new features and enhancements to the Qualcomm AI Stack SDK.

  • Participate in design reviews and code reviews, ensuring software quality and maintainability.

  • Mentor junior engineers helping them prioritize work, and drive execution across multiple initiatives.

Requirements

Do you have experience in Tooling?, This role requires strong technical ownership, the ability to work independently, and the capability to drive features end-to-end while mentoring junior engineers., * Bachelor’s degree in Computer Science, Engineering, Information Systems, or related field and 4+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience., Master’s degree in Computer Science, Engineering, Information Systems, or related field and 3+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience. OR PhD in Computer Science, Engineering, Information Systems, or related field and 2+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience., * Bachelor’s degree in computer science, computer engineering, or a related field and 6+ years (Staff) / 8+ years (Sr. Staff) of experience in software design, development, and delivery.

  • OR

  • Master’s degree/PhD in computer science, computer engineering, or a related field and 5+ years (Staff) / 7+ years (Sr. Staff) of experience in software design, development and delivery.

  • 3+ years of hands-on experience in AI/ML software development, with a focus on inference or model optimization.

  • Strong understanding of AI/ML fundamentals, including deep learning and inference pipelines.

  • Deep understanding of transformer architectures, attention mechanisms, and performance tradeoffs.

  • Proficiency in Python and C/C++ for production-quality software development.

  • Experience working with PyTorch and ONNX models and tooling.

  • Debugging skill of complex issues, perform root cause analysis, and ensure high system reliability

  • Ability to work independently, collaborate across teams, and drive complex features end-to-end.

Preferred Qualifications

  • Working knowledge of graph theory, graph optimizations, and compiler-style transformations.

  • Experience with LLM, LVM, and LMM inference pipelines, including prefill and generation workflows.

  • Familiarity with Hugging Face ecosystem, including model repositories and interfaces such as PEFT.

  • Experience with LoRA, MoE-based models, and awareness of modern GenAI inference techniques.

  • Experience with Android and/or RTOS environments (e.g., QNX).

  • Experience with CMake-based build environments, agile software development practices, and git-based SCM.

  • 2+ years of experience in embedded software or system-level software development and optimization.

  • At least 2 years of experience interacting with senior leadership (Director level and above).

  • Ability to collaborate across a globally diverse team and manage multiple priorities.

  • Previous experience of mentoring junior engineers

  • User-level or development experience with Qualcomm AI Stack / SDKs (e.g., QAIRT, QNN, Genie).

  • Exposure to Snapdragon SoCs and AI accelerators such as NPU.

  • Prior hands-on experience with GenAI features such as transformer architectures, LoRA, MoE, speculative decoding, and vision encoder/decoder models.

Benefits & conditions

4.04.0 out of 5 stars Raleigh, NC $158,400 - $237,600 a year - Full-time, Qualcomm expects its employees to abide by all applicable policies and procedures, including but not limited to security and other requirements regarding protection of Company confidential information and other confidential and/or proprietary information, to the extent those requirements are permissible under applicable law.

Pay range and Other Compensation & Benefits : $158,400.00 - $237,600.00

The above pay scale reflects the broad, minimum to maximum, pay scale for this job code for the location for which it has been posted. Even more importantly, please note that salary is only one component of total compensation at Qualcomm. We also offer a competitive annual discretionary bonus program and opportunity for annual RSU grants (employees on sales-incentive plans are not eligible for our annual bonus). In addition, our highly competitive benefits package is designed to support your success at work, at home, and at play. Your recruiter will be happy to discuss all that Qualcomm has to offer - and you can review more details about our US benefits at this link .

About the company

Qualcomm Technologies, Inc., As a leading technology innovator, Qualcomm pushes the boundaries of what’s possible to enable next-generation experiences and drive digital transformation, creating a smarter, connected future for all.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · WWC 2023

1:15 min

Defining classes and building packages with pybind11

Konstantin Bespalov · WWC 2023

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

4:41 min

Replacing PyTorch with ONNX runtime for AWS Lambda deployments

Marek Suppa · LIVE

1:49 min

Augmenting junior and principal engineering roles with AI

Neel Sundaresan Neel Sundaresan +1 · WWC Europe 2026

Videos

See all

Related articles

See all