Software Engineer, ML platform and Infrastructure

Apple Inc.
Austin, TX, United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Compensation
$212,000.0 - $318,400.0
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Application Programming Interfaces (APIs) Amazon Web Services Systems Engineering DevOps Distributed Data Store Distributed Systems Python (Programming Language) Machine Learning Network Protocols Open Source Technology Reliability Engineering
+13 more
Azure Machine Learning Software Engineering Data Processing Graphics Processing Unit (GPU) Large Language Models Snowflake Multi-Agent Systems Apache Spark Generative AI Backend Kubernetes Apache Flink Machine Learning Operations

Job description

Join Apple’s Applied Machine Learning Team as a Machine Learning Platform Engineer and play a central role in designing and building the systems that power our Data, Machine Learning, and Generative AI initiatives. You will architect and engineer robust, high-performance, massively scalable platforms that serve as the foundation for groundbreaking ML workloads across the enterprise.

Requirements

Do you have experience in Software engineering?, We are looking for talented Software Engineers who are passionate about distributed systems and large-scale infrastructure to build and operate world-class ML platforms and products across cloud environments., In this role, you will apply software engineering depth to solve the hardest challenges in large-scale distributed systems-designing for reliability, performance, and efficiency from the ground up. You will own the technical direction of ML/Data/Inference platform capabilities, leading the evaluation and integration of cutting-edge open-source technologies and building innovative internal solutions that raise the bar for scalability and resilience across our ML ecosystem. You’ll collaborate closely with cross-functional engineering and business teams, influencing technical strategy and contributing meaningfully to the broader platform roadmap.”,”responsibilities”:”Highly proficient in Python, Java, or Go, with a strong track record of building production-grade automation, tooling, and system-level software.

Deep understanding of LLM infrastructure requirements-including GPUs, TPUs, and Inferentia-with hands-on experience engineering systems that optimize their utilization and performance.

Experience designing and building Agents and MCP servers, with hands-on expertise in frameworks such as LangGraph and LangChain.

Solid background in software engineering for complex, large-scale distributed systems, with strong familiarity with DevOps and reliability engineering practices.

Expert-level proficiency with AWS/GCP and deep, hands-on experience architecting and engineering containerized workloads using Kubernetes in production environments.

Proven ability to read, understand, and make meaningful contributions to complex open-source codebases in the ML infrastructure space.

Strong command of operating system internals, networking protocols, and security principles, applied to building highly available and resilient systems.

Exceptional analytical and problem-solving skills, with a demonstrated ability to identify and resolve critical system bottlenecks and failures in high-stakes environments.

Preferred Qualifications

Experience engineering scalable solutions for data processing and model training/fine-tuning workflows.

Hands-on experience building with distributed data technologies for ML training such as Spark, Flink, Iceberg, or Snowflake, with a deep understanding of their architectural trade-offs at scale.

Minimum Qualifications

5+ years of experience in software development, with a strong focus on backend systems and APIs.

2+ years of experience working with LLMs, Agent Frameworks

5+ years of experience with cloud platforms such as AWS,or GCP

Benefits & conditions

4.14.1 out of 5 stars Austin, TX $212,000 - $318,400 a year, Pulled from the full job description

  • Employee stock purchase plan
  • Health insurance
  • Retirement plan
  • Dental insurance
  • RSU

About the company

The Applied Machine Learning team has been at the forefront of accelerating digital transformation through machine learning across Apple’s enterprise ecosystem. Our ML Platforms, Solutions, and Services deliver a comprehensive suite of capabilities that drive efficiency, agility, and innovation at Apple scale-serving business-critical needs across the enterprise.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

1:33 min

Integrating internal APIs and maintaining data sovereignty

Mahran Meißner Mahran Meißner · WWC Europe 2026

2:14 min

Solving complex platform architecture challenges at an enterprise scale

Maria Apazoglou · Coffee With Developers

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

Videos

See all

Related articles

See all