On-Device ML Infrastructure Engineer (ML User Experience APIs), Graphics, Games and Machine Learning

Apple Inc.
Cupertino, CA, United States
2 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) C++ (Programming Language) Computer Engineering Python (Programming Language) Machine Learning Pytorch Information Technology HuggingFace Machine Learning Operations Api Design

Job description

Experteer Overview In this role you will help shape Apple’s ML infrastructure by building model conversion and authoring APIs for end-to-end model deployment on Apple platforms. You will demonstrate and optimize cross-ecosystem workflows, integrating models from external repositories like Hugging Face with Apple’s stack. You’ll design and stress-test optimizations across source-level code and Apple representations to achieve strong performance on diverse hardware. This is a hands-on, impact-driven position at the intersection of research, software, and hardware engineering, focused on delivering a premier on-device ML experience. Compensation / Benefits * Develop and expose ML model conversion and authoring APIs as the main entry point into Apple’s ML infrastructure * Onboard popular ML models via end-to-end workflows highlighting authoring and runtime capabilities * Integrate Apple ML tools into internal and external model repositories (e.g., Hugging Face) * Ideate, design and stress test optimizations from PyTorch programs to custom transformations in Apple’s model representation * Support engineering of end-to-end inference stack from model creation to deployment Tasks * Confirmed understanding of ML modeling (architectures, training vs. inference trade-offs) * Experience in ML deployment optimizations (quantization) * Strong Python API design experience * Proficiency in Python and familiarity with C++ * Experience with ML authoring frameworks (e.g., PyTorch, MLX, JAX) * Experience with MLIR/LLVM or similar compiler toolchains * Familiarity with Hugging Face or other model repositories * Bachelor in Computer Science, Engineering, or related subject area * Hands-on experience with ML inference optimizations (quantization, pruning, KV caching) * Strong communication skills and ability to engage multi-functional audiences Key requirements *

Requirements

Experteer Overview In this role you will help shape Apple’s ML infrastructure by building model conversion and authoring APIs for end-to-end model deployment on Apple platforms. You will demonstrate and optimize cross-ecosystem workflows, integrating models from external repositories like Hugging Face with Apple’s stack. You’ll design and stress-test optimizations across source-level code and Apple representations to achieve strong performance on diverse hardware. This is a hands-on, impact-driven position at the intersection of research, software, and hardware engineering, focused on delivering a premier on-device ML experience. Compensation / Benefits * Develop and expose ML model conversion and authoring APIs as the main entry point into Apple’s ML infrastructure * Onboard popular ML models via end-to-end workflows highlighting authoring and runtime capabilities * Integrate Apple ML tools into internal and external model repositories (e.g., Hugging Face) * Ideate, design and stress test optimizations from PyTorch programs to custom transformations in Apple’s model representation * Support engineering of end-to-end inference stack from model creation to deployment Tasks * Confirmed understanding of ML modeling (architectures, training vs. inference trade-offs) * Experience in ML deployment optimizations (quantization) * Strong Python API design experience * Proficiency in Python and familiarity with C++ * Experience with ML authoring frameworks (e.g., PyTorch, MLX, JAX) * Experience with MLIR/LLVM or similar compiler toolchains * Familiarity with Hugging Face or other model repositories * Bachelor in Computer Science, Engineering, or related subject area * Hands-on experience with ML inference optimizations (quantization, pruning, KV caching) * Strong communication skills and ability to engage multi-functional audiences Key requirements *

About the company

On-Device ML Infrastructure Engineer (ML User Experience APIs), Graphics, Games and Machine Learning

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:10 min

Understanding the core concepts of API design

Alen Pokos · LIVE

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · WWC 2023

1:12 min

Training and fine-tuning models natively using MLX

MIlan Todorović MIlan Todorović · WWC 2025

3:26 min

Prioritizing backward compatibility in API design

Justin Kitagawa · Coffee With Developers

4:41 min

Replacing PyTorch with ONNX runtime for AWS Lambda deployments

Marek Suppa · LIVE

40 sec

Navigating Apple's evolving on-device AI and machine learning stack

Precious Osaro Precious Osaro · WWC Europe 2026

Videos

See all

Related articles

See all