Sr. AI Inference Platform Engineer

Apple Inc.
Seattle, WA, United States
22 days ago
Apply on www.seattlejobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Data Analysis C++ (Programming Language) Continuous Integration Distributed Systems Python (Programming Language) Prometheus Software Engineering AI Infrastructure Data Logging Large Language Models Grafana
+8 more
Kubernetes Information Technology Data Analytics TensorRT Splunk Data Pipelines Golang Programming Languages

Job description

We are looking for senior engineer to build tooling, automation, and analysis capabilities that strengthen our AI inference platform. This role will focus on developing sophisticated performance benchmarking systems, capacity projection models, and data analysis pipelines that directly inform our AI infrastructure teams and capacity planners. You’ll work at the intersection of AI systems performance, distributed infrastructure, and software engineering to help the team make data-driven decisions about scaling and optimizing our inference platform.

Requirements

  • BS or MS in Computer Science or related technical field.
  • Solid understanding of AI/ML inference architecture and the performance characteristics of serving systems.
  • 7 or more years of experience with performance and infrastructure engineering in distributed systems.
  • 7 years of experience coding in Python, Go, C++, or other programming languages.
  • Experience with automation engineering, tooling, and data pipelines to support engineering workflows.
  • Strong knowledge of GPU/accelerator architecture as it relates to AI workloads.
  • Practical statistical knowledge applicable to performance analysis and forecasting.
  • Excellent communication skills and ability to turn data into clear guidance for infrastructure teams and capacity planners., * Experience with performance benchmarking and methodologies for AI/ML inference systems.
  • Familiarity with capacity planning and forecasting/projection models for large-scale infrastructure.
  • Experience with GPU profiling and observability tools (e.g., Nsight, other vendor-specific profilers).
  • Experience with data visualization and reporting tools/frameworks for surfacing performance trends to stakeholders.
  • Familiarity with ML serving frameworks and runtimes (e.g., Triton, TensorRT-LLM, vLLM, or similar).
  • Experience with CI/CD and workflow orchestration tools for building automated performance analysis pipelines.
  • Knowledge of cluster schedulers and orchestration platforms (e.g., Kubernetes).
  • Experience with metrics and logging tools (e.g., Prometheus, Grafana, Splunk).

About the company

At Apple, we believe the future of AI is defined not just by models, but by the infrastructure that powers them. Our AI inference platform sits at the heart of products and experiences used by hundreds of millions of people worldwide, and we are building the systems that ensure it scales reliably, efficiently, and intelligently.

As part of our next-generation datacenter engineering team, you will play a critical role in shaping how we understand, measure, and grow our AI infrastructure. You will design and build the tooling and analysis systems that give our engineers and capacity planners a clear, real-time picture of performance across our fleet. Your work will directly influence how we invest in hardware, how we detect regressions before they reach production, and how we forecast capacity needs months in advance.

This is a high-impact, cross-functional role for an engineer who is energized by complexity, thrives on turning raw data into actionable insight, and wants to work on problems that matter at massive scale.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.seattlejobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · World Congress 2025

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

Videos

See all

Related articles

See all