AI Researcher

Qualcomm
San Diego, CA, United States
4 days ago
Apply on find.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$200,000.0 - $250,000.0
Working hours
Regular working hours
Job source

Tech stack

C++ (Programming Language) Distributed Computing Environment Python (Programming Language) Delivery Pipeline Large Language Models Deep Learning Low Latency Machine Learning Operations

Job description

Qualcomm seeks an AI Researcher focused on on-device LLM efficiency to advance next generation wireless and mobile platforms. You will research and prototype algorithms for model compression, quantization, pruning, distillation, and on device inference optimization across smartphones, IoT, and automotive systems. Collaborating with cross functional silicon, software, and product teams, you will design experiments, benchmark models, and translate research into production ready solutions. This role offers broad impact on 5G and edge AI, strong publication opportunities, and growth in a fast paced, innovation driven environment., * Research and develop methods for on-device LLM efficiency, including compression, quantization, pruning, and distillation

  • Prototype and benchmark LLM inference on Qualcomm mobile, Io
  • T, and automotive platforms
  • Collaborate with silicon, software, and product teams to translate research into deployable solutions
  • Design and run experiments to evaluate accuracy, latency, power, and memory trade-offs
  • Publish and present findings internally and externally to help shape Qualcomm’s AI roadmap
  • Contribute to tooling and pipelines for efficient model deployment on edge devices

Requirements

  • Large language models (LLMs)
  • Model compression and quantization
  • Network pruning and distillation
  • On-device and edge AI optimization
  • Deep learning frameworks (Py
  • Torch, Tensor
  • Flow)
  • C++ and Python programming
  • GPU/NPUs and hardware-aware MLPerformance profiling and benchmarking
  • Distributed training and experimentation
  • MLOps and deployment pipelines

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on find.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:42 min

Dissecting artificial intelligence layers from compute to applications

Christian Nagel Christian Nagel +3 · World Congress 2026 Europe

5:48 min

Balancing delivery latency with stream reliability and scale

Phil Cluff · LIVE

1:25 min

Distinguishing artificial intelligence from deep learning

Sam Witteveen · Coffee With Developers

1:51 min

Unifying software compliance into standard delivery pipelines

Marcus Ross Marcus Ross · World Congress 2026 Europe

2:36 min

Exploring high-level Python frameworks for accelerated enterprise artificial intelligence

Paul Graham Paul Graham · LIVE

3:37 min

Accessing API documentation and testing remote driving latency

Alexandru Ciinaru Alexandru Ciinaru +3 · World Congress 2025

Videos

See all

Related articles

See all