Research Engineer (AI/RL Infrastructure)

INC Research
Sunnyvale, CA, United States
3 months ago
Apply on welcometothejungle.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$126,000.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Big Data Computer Clusters Nvidia CUDA Software Debugging Machine Learning Open Source Technology Robotic Automation Software Pytorch Build Management Kubernetes Machine Learning Operations
+1 more
Data Pipelines

Job description

  • We are looking for a passionate Research Engineer (AI/RL Infrastructure) to join the Research Group at Applied Intuition
  • This role is ideal for engineers who design, build, and operate state-of-the-art, large-scale ML systems and enjoy working closely with researchers to develop and accelerate the core platform powering next-generation physical AI systems
  • The mission of the Research Group is to create cutting-edge technology enabling next-generation physical AI, with emphasis on the two most challenging applications reshaping our everyday life: end-to-end autonomous driving and robotic generalist
  • We have a group composed of leading experts from top institutions and companies, recognized for their exceptional academic and industry contributions-including eight Best Paper awards at premier conferences and journals such as CVPR and ICRA
  • Learn more at appliedintuition.com/research
  • Supported by industry-leading tools and infra, researchers can access millions of miles of data from large fleets, and deploy methods they develop into various autonomous and robotic systems including self-driving cars/trucks, autonomous mining/construction machines, humanoid robots and dexterous hands
  • In addition to your research contributions, you will contribute to and learn from best practices in the autonomy and robotics industries within our fast-paced and customer-focused culture
  • Improvements deployed to our system immediately help our customers with their programs and deliver value to our business
  • Design and build training and evaluation infrastructure to support our current AI research directions, orchestrating massive GPU clusters to process PBs of multimodal sensor data
  • Build robust benchmarking, continuous evaluation, and regression tracking systems to measure model performance across diverse, long-tail real-world driving distributions
  • Develop large-scale data sampling, dataset generation, and advanced data curation pipelines, leveraging state-of-the-art AI models to power a closed-loop data flywheel
  • Enable high-throughput distributed training across heterogeneous cloud environments, focusing on reliability, efficiency, and cost-aware scaling
  • Collaborate closely with AI research, autonomy, and platform teams to translate cutting-edge research into production-ready systems

Requirements

  • Experience building and operating production-grade software systems across the full machine learning lifecycle, including training, evaluation, data, and deployment
  • Experience with performance engineering and compute acceleration for large-scale ML training, including profiling, bottleneck analysis, and optimization
  • Strong systems-level debugging skills to diagnose and resolve issues in large-scale distributed training, spanning model code, data pipelines, runtimes, and cluster infrastructure
  • Opinions about building a company-wide platform for ML training, evaluation, and deployment
  • Technical experience in: Pytorch, CUDA, Ray, Flyte, K8s
  • Deep familiarity with the open-source ML and systems ecosystem, with judgment on when to adopt open source versus build in-house
  • Industry experience on relevant topics (self-driving application preferred)
  • Don’t meet every single requirement? If you’re excited about this role but your past experience doesn’t align perfectly with every qualification in the job description, we encourage you to apply anyway. You may be just the right candidate for this or other roles

Benefits & conditions

  • Health Insurance: Applied covers 100% of the cost and offers medical, dental, vision, and life insurance
  • Fitness Stipend: Our wellness program supports your health and fitness goals. We pay for your trainers, gym memberships, massages, and more!
  • 401(k) Match: 50% employer matching to help you save money for retirement
  • Learning Stipend: We invest in your learning and growth by reimbursing you for educational and training expenses
  • Parental leave: 12-week, fully paid parental leave to bond with a newborn child or new adopted or foster child
  • Catered lunch, expensed dinner, on-site meals and snacks

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on welcometothejungle.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · World Congress 2022

1:06 min

Compiling PyTorch environments for advanced time forecasting

Christoph Lohrmann Christoph Lohrmann +1 · World Congress 2026 Europe

2:10 min

Why organizations combine big data and machine learning

Ayon Roy · LIVE

Videos

See all

Related articles

See all