machine learning and GPU programming engineers

Apple Inc.
United States
about 2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

C++ (Programming Language) Compilers Nvidia CUDA Computer Programming Data Centers Distributed Computing Environment Linux Kernel Machine Learning Objective-C (Programming Language) Tensorflow Private Cloud Environment Graphics Processing Unit (GPU)
+6 more
Large Language Models Generative AI Gpu Programming Hardware Infrastructure Stable Diffusion Software Library

Requirements

Apple’s Server ML Frameworks team in GPU, Graphics and Machine Learning works on enabling Apple Intelligence through high-performance, distributed inference of GenAI applications (such as LLMs) on Private Cloud Compute. You will get to work on custom-built server hardware that brings the power and security of Apple silicon to the data center. We are looking for engineers with systems background who are deeply passionate about building scalable, efficient, and production-grade solutions tailored for high-throughput GPU execution.

Our team is seeking extraordinary machine learning and GPU programming engineers who are passionate about providing robust compute solutions for accelerating Machine learning libraries on Apple Silicon. Role has the opportunity to influence the design of compute and programming models in next generation GPU architectures.

3+ years of programming and problem-solving experience with C/C++/ObjC\nExperience with GPU kernel development & optimizations using compute programming models such as Metal, CUDA etc.\nExperience with Distributed training or inference techniques\nExperience with system level programming and computer architecture

Experience with graph compilers such as CuTE, CuTile, Triton, OpenXLA or LLVM is a plus\nGood understanding of LLM and Diffusion based model architectures

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:12 min

Training and fine-tuning models natively using MLX

MIlan Todorović MIlan Todorović · WWC 2025

6:21 min

Previewing upcoming hardware acceleration capabilities for Python environments

Chris Heilmann +2 · LIVE

1:39 min

Fundamentals of tensors and the TensorFlow library

Håkan Silfvernagel · LIVE

1:51 min

Evolution of custom compilers and virtual machines

Florian Rappl · LIVE

1:37 min

Accelerating compute with focused developer tools

Julia Koch Julia Koch +1 · WWC Europe 2026

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · WWC Europe 2026

Videos

See all

Related articles

See all