> Markdown version of [/jobs/ext/3095732-ml-data-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/3095732-ml-data-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # ML Data Infrastructure Engineer - **Company:** AppLovin Corporation - **Location:** Palo Alto, United States - **Experience:** Starter - **Salary:** $150,000.0 - $224,000.0 - **Contract:** Permanent contract - **Skills:** Training Data, Data Infrastructure, Data Structures, Distributed Computing Environment, Distributed Systems, Fault Tolerance, Machine Learning, Performance Tuning, Raw Data, Software Systems, Data Processing, System Availability, Apache Spark, Build Management, Information Technology, Apache Flink, Machine Learning Operations - **Published:** September 26, 2026 - **Apply:** https://www.dice.com/job-detail/d77b9414-f489-427f-93c9-ac105cb7504b ## About the Role * Have 1 - 3 years of experience and a minimum of a BS and/or MS in Computer Science * Strong software engineering fundamentals, with experience building high-throughput, fault-tolerant distributed systems * Hands-on experience with distributed computing frameworks such as Apache Spark or Flink * Solid grounding in data structures, systems design, and performance optimization * Strong problem-solving skills and attention to detail Preferred Qualifications * Background in MLOps, Data Infrastructure, or ML Infrastructure * Experience with ML training pipelines, feature stores, or model-serving systems ## Description As a member of our ML Data Platform team, you'll solve technical challenges, including upgrading and implementing state-of-the-art software infrastructure. The team builds a high-performance, high availability, globally distributed ecosystem platform of services that in turn provide the foundation for rapid development of novel new systems that integrate into that ecosystem and improve it. The Impact You'll Make * Design and build data processing infrastructure for model training and feature serving, optimizing for performance, reproducibility, and traceability * Collaborate closely with research teams to design and implement novel data processing architectures for emerging model and training paradigms * Identify and resolve performance bottlenecks across the training data pipeline, from raw data ingestion to feature delivery * Establish best practices, tooling for data infrastructure used across ML teams