AIML - Sr Data Engineer, DMLI

Apple Inc.
Cupertino, CA, United States
12 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours

Tech stack

Training Data Artificial Intelligence Airflow Data Analysis Automated Storage and Retrieval Systems BigQuery Cloud Engineering Code Review Encodings Continuous Integration Information Engineering Data Governance
+30 more
Data Infrastructure Data Integration Extract Transform Load (ETL) Dataspaces Data Systems Software Debugging Distributed Systems Data Flow Control Python (Programming Language) Machine Learning Meta-Data Management Performance Tuning Query Optimization Cloudera Software Engineering Data Streaming Google Cloud Feature Engineering Large Language Models Apache Spark Generative AI Event Driven Architecture Information Technology Data Lineage Data Management Machine Learning Operations Artificial Intelligence Markup Language (AIML) Software Version Control Data Pipelines Apache Beam

Job description

Are you excited to tackle some of the most ambitious technical challenges in Apple Intelligence? Be involved in collaborating closely with our machine learning researchers, engineers, and data scientists? Together, you will orchestrate groundbreaking research initiatives and develop transformative products designed to build a significant impact for billions of users worldwide!, The AI and Machine Learning team is looking for Senior Data Engineer to build world-class data infrastructure and solutions that are used by data scientists, ML engineers and researchers to power Apple Foundation Model lifecycle., We are looking for a skilled Senior Data Engineer to join our engineering team. In this role, you will be responsible for designing, building, and maintaining scalable data pipelines, data platforms, and data integration solutions that enable reliable analytics, machine learning, and business intelligence capabilities across the organization.

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical field (or equivalent practical experience).
  • Strong software engineering foundation with 5+ years of experience in data engineering, backend engineering, or distributed systems.
  • Strong hands-on experience designing and building large-scale batch and streaming data pipelines using Python and distributed systems.
  • Expertise in ETL/ELT development, event-driven architectures, and working with high-volume datasets using technologies such as Pub/Sub, Dataflow (Apache Beam), and Spark-based processing (Dataproc or equivalent).
  • Deep expertise in the Google Cloud Platform data ecosystem, including BigQuery, Dataflow, Pub/Sub, and Composer (Airflow). Proven ability to design, optimize, and operate scalable data platforms, with strong experience in BigQuery performance tuning, cost optimization, data modeling, partitioning, clustering, and query optimization at scale.
  • Proficient in Python, version control, CI/CD practices, testing, and code reviews., * 7+ years of experience in data engineering or large-scale distributed systems.
  • Experience designing end-to-end data platforms or data mesh architectures.
  • Strong understanding of data governance, data lineage, and metadata management frameworks.
  • Experience supporting ML pipelines and MLOps workflows, including feature engineering and training data generation.
  • Experience building systems with strong reliability, observability, and SLAs (monitoring, alerting, debugging distributed pipelines).
  • Familiarity with Generative AI / LLM-based systems, including: LLM-powered data workflows, Agentic pipelines, Embedding/vector-based retrieval systems.
  • Experience influencing architecture decisions across teams and driving technical direction.
  • Experience collaborating in cross-functional, fast-paced tech environments or cloud-native organizations with a focus on building reliable, maintainable, and production-grade data systems.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.techcareers.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · WWC Europe 2026

3:27 min

Explaining query execution overhead and caching limitations in BigQuery

Adnan Rahic · JS Congress

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:05 min

Audience questions on AI agents and pipeline vectorization

Joy Joy · WWC 2024

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

2:30 min

Leveraging BigQuery ML for scalable SQL-based segmentation experiments

Julian Joseph · LIVE

Videos

See all

Related articles

See all