Senior Data Engineer, AI Systems

Movable Ink
New York, NY, United States
about 1 month ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$165,000.0 - $215,000.0
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Unit Testing Big Data Code Review Continuous Delivery Continuous Integration Information Engineering Data Infrastructure Distributed Systems Github Python (Programming Language)
+20 more
Machine Learning Performance Tuning Query Optimization Recommender Systems Cloudera Software Engineering Parquet Google Cloud Data Storage Technologies Cloud Platform System Apache Spark Git Data Layers Data Lakes Pyspark Kubernetes Deployment Automation Apache Kafka Data Pipelines Docker

Job description

The AI Systems team owns the core recommendations engine and ML platform that powers billions of AI-driven marketing decisions daily across some of the world’s largest consumer brands. As a Senior Data Engineer, you will own the Spark-based data pipelines and data infrastructure at the heart of this system - building, scaling, and optimizing the data layer that feeds our production ML models. You will work alongside ML engineers and scientists in a collaborative environment, contributing data pipelines and products to power our core recommender systems and our DaVinci Personalization product. This is an opportunity to work end-to-end on large-scale data systems that touch millions of customers, on a team working at the intersection of data engineering and machine learning.

This role will be reporting to the Director of Engineering (AI/ML)., * Build, maintain, and optimize production data pipelines that power AI-driven personalization at scale across content selection, send-time optimization, subject line personalization, and frequency capping

  • Own and scale Spark-based batch pipelines, including cluster configuration, tuning, and performance optimization across GCP Dataproc
  • Build and maintain our ML Data Lake, ensuring data quality, accessibility, and efficient storage
  • Support the data needs of ML Engineers and Scientists for model development, training, and evaluation
  • Identify and resolve performance bottlenecks and scaling limitations in data pipelines and infrastructure
  • Collaborate with distributed systems engineers on the platform’s architectural evolution, ensuring data layer continuity throughout
  • Continuously improve data infrastructure for greater scalability and reliability
  • Release features and data products that deliver measurable and tangible business value

Requirements

  • 5+ years of data engineering experience
  • Deep expertise with Apache Spark, including the PySpark DataFrame API and experience solving challenging scaling problems
  • Experience with large-scale data processing, cluster configuration, optimization, and tuning (we use GCP Dataproc)
  • Strong software development skills in Python (unit testing, git, code review, CI/CD)
  • Experience with data storage formats (we use Parquet, Delta Lake)
  • Experience with event streaming data (we use Kafka)
  • Experience with cloud computing platforms (we use Google Cloud Platform)
  • Experience with advanced query optimization
  • Familiar with Software Development Lifecycle practices, such as continuous integration/continuous delivery and automated deployment (we use Docker, Kubernetes, and GitHub Actions)
  • Ability to collaborate with technical partners - you’ll be working closely with ML engineers, scientists, and other teams to determine requirements and make design decisions
  • Enjoys working in a fast-paced, goal-driven environment

Benefits & conditions

The base pay range for this position is $165K - $215K USD/year, which can include additional bonus depending on the position ultimately offered, in addition to a full range of medical, financial, and/or other benefits. The base pay offered may vary depending on job-related knowledge, skills, and experience.

About the company

Movable Ink scales content personalization for marketers through data-activated content generation and AI decisioning. The world’s most innovative brands rely on Movable Ink to maximize revenue, simplify workflow and boost marketing agility. Headquartered in New York City with close to 600 employees, Movable Ink serves its global client base with operations throughout North America, Central America, Europe, Australia, and Japan.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

Videos

See all

Related articles

See all