> Markdown version of [/jobs/ext/176120-senior-research-engineer-training-data-infrastructure-in-foundation-models](https://www.wearedevelopers.com/jobs/ext/176120-senior-research-engineer-training-data-infrastructure-in-foundation-models). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Research Engineer, Training Data Infrastructure in Foundation Models - **Company:** Apple Inc. - **Location:** Cupertino, CA, United States - **Experience:** Expert - **Salary:** $181,100.0 - $318,400.0 - **Contract:** Permanent contract - **Skills:** Training Data, Artificial Intelligence, Amazon Web Services, Big Data, C++ (Programming Language), Cloud Computing, Cloud Engineering, Profiling, Data Deduplication, Data Governance, Data Infrastructure, Data Systems, Software Debugging, File Systems, Distributed Computing Environment, Distributed Systems, Python (Programming Language), Machine Learning, Performance Tuning, Software Configuration Management, Software Engineering, Systems Architecture, Data Processing, Data Storage Management, Data Storage Technologies, Large Language Models, Backend, Information Technology, Build Tools, Data Pipelines, Apache Beam, Data Generation - **Published:** May 27, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=f9d963716defe7cb ## About the Role Do you have experience in System performance optimization?, Research Collaboration: Experience working within or closely with ML research organizations (e.g., as a Research Engineer), with an ability to translate research results into engineering implementations. Domain Knowledge: Familiarity with lifecycle of modern LLM training, end-to-end workflows, and underlying system architecture. Complex Data Types: Experience in processing complex data modalities beyond plain text, such as source code repositories, images, videos, and audios. Minimum Qualifications Education: Bachelor's degree in Computer Science, Electrical Engineering, or Mathematics. Technical Expertise: 4+ years of software engineering experience with a specific focus on Data Infrastructure, Distributed Systems, or AI/ML Engineering. Language Proficiency: Expert fluency in Python, and strong competence in system languages such as C++. Cloud Architecture: Extensive experience architecting solutions on major public cloud platforms (e.g. GCP) to build scalable data systems (e.g. with Apache Beam, GCS) Performance Engineering: Deep experience profiling and optimizing high-throughput data systems. Demonstrated ability to debug distributed bottlenecks (e.g., stragglers, I/O saturation), optimize data formats and provide efficient data storage solutions. ## Description This position operates at the convergence of Software Engineering and Machine Learning Research. Unlike traditional backend roles, this position requires you to design systems where the outcome is the statistical distribution and quality of data itself. You will work alongside Research Scientists to transform theoretical observations into concrete, scalable engineering solutions. Your core focus will be the architecture of our Data Acquisition, Processing, and Repository Management systems for Large Model training. You will lead technical efforts to enable active, quality-driven data curation, including filtering, deduping, synthetic data generation and data mixing, ensuring our models are trained on the highest-quality information available.","responsibilities":"Architect Scalable Ingestion Systems: Design and implement high-throughput distributed systems to ingest petabytes of text and multimodal data from diverse sources, including web crawls and third-party partnerships. Repository Optimization: Manage the lifecycle of large-scale datasets across data storage and high-performance file systems. Optimize data formats for efficient random access and sequential scanning during model training. Data Governance & Privacy: Engineer robust data governance and privacy solutions for the training data, in collaboration with compliance and legal teams, to ensure adherence to stringent regulatory standards. High-Performance Processing Pipelines: Build and maintain distributed data processing workflows using advanced frameworks on cloud infrastructure (e.g., GCP, AWS). Algorithmic Data Curation: Implement sophisticated data filtering and selection logic to remove low-quality content. Develop semantic deduplication at scale to prevent model memorization and improve training efficiency. Decontamination Removal: Design automated systems to detect and remove benchmark leakage, ensuring that evaluation datasets remain strictly isolated from training corpora. Infrastructure for Scaling Laws: Collaborate with researchers to enable data ablations and scaling experiments. Build tools to support systematic data mixture optimization and empirically data studies. ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Profiling Symfony & PHP apps with Blackfire](https://www.wearedevelopers.com/videos/265-profiling-symfony-php-apps-with-blackfire) - [Developer Experience, Platform Engineering and AI powered Apps](https://www.wearedevelopers.com/videos/990-developer-experience-platform-engineering-and-ai-powered-apps) - [Nest.js - TypeScript in the backend can also be clean](https://www.wearedevelopers.com/videos/1033-nest-js-typescript-in-the-backend-can-also-be-clean) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)