> Markdown version of [/jobs/ext/2547035-senior-data-engineer-ai-systems](https://www.wearedevelopers.com/jobs/ext/2547035-senior-data-engineer-ai-systems). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Data Engineer, AI Systems - **Company:** Movable Ink - **Location:** New York, NY, United States (Remote available) - **Experience:** Expert - **Salary:** $165,000.0 - $215,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Unit Testing, Big Data, Code Review, Continuous Delivery, Continuous Integration, Information Engineering, Data Infrastructure, Distributed Systems, Github, Python (Programming Language), Machine Learning, Performance Tuning, Query Optimization, Recommender Systems, Cloudera, Software Engineering, Parquet, Google Cloud, Data Storage Technologies, Cloud Platform System, Apache Spark, Git, Data Layers, Data Lakes, Pyspark, Kubernetes, Deployment Automation, Apache Kafka, Data Pipelines, Docker - **Published:** August 3, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=f7ee9658ab84208e ## About the Role * 5+ years of data engineering experience * Deep expertise with Apache Spark, including the PySpark DataFrame API and experience solving challenging scaling problems * Experience with large-scale data processing, cluster configuration, optimization, and tuning (we use GCP Dataproc) * Strong software development skills in Python (unit testing, git, code review, CI/CD) * Experience with data storage formats (we use Parquet, Delta Lake) * Experience with event streaming data (we use Kafka) * Experience with cloud computing platforms (we use Google Cloud Platform) * Experience with advanced query optimization * Familiar with Software Development Lifecycle practices, such as continuous integration/continuous delivery and automated deployment (we use Docker, Kubernetes, and GitHub Actions) * Ability to collaborate with technical partners - you'll be working closely with ML engineers, scientists, and other teams to determine requirements and make design decisions * Enjoys working in a fast-paced, goal-driven environment ## Description The AI Systems team owns the core recommendations engine and ML platform that powers billions of AI-driven marketing decisions daily across some of the world's largest consumer brands. As a Senior Data Engineer, you will own the Spark-based data pipelines and data infrastructure at the heart of this system - building, scaling, and optimizing the data layer that feeds our production ML models. You will work alongside ML engineers and scientists in a collaborative environment, contributing data pipelines and products to power our core recommender systems and our DaVinci Personalization product. This is an opportunity to work end-to-end on large-scale data systems that touch millions of customers, on a team working at the intersection of data engineering and machine learning. This role will be reporting to the Director of Engineering (AI/ML)., * Build, maintain, and optimize production data pipelines that power AI-driven personalization at scale across content selection, send-time optimization, subject line personalization, and frequency capping * Own and scale Spark-based batch pipelines, including cluster configuration, tuning, and performance optimization across GCP Dataproc * Build and maintain our ML Data Lake, ensuring data quality, accessibility, and efficient storage * Support the data needs of ML Engineers and Scientists for model development, training, and evaluation * Identify and resolve performance bottlenecks and scaling limitations in data pipelines and infrastructure * Collaborate with distributed systems engineers on the platform's architectural evolution, ensuring data layer continuity throughout * Continuously improve data infrastructure for greater scalability and reliability * Release features and data products that deliver measurable and tangible business value ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)