> Markdown version of [/jobs/ext/687312-data-infra-engineer](https://www.wearedevelopers.com/jobs/ext/687312-data-infra-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Infra Engineer - **Company:** LGA AIRPORT RESTAURANTS, L.P. - **Location:** United States - **Experience:** Expert - **Salary:** $175,000.0 - $221,000.0 - **Contract:** Permanent contract - **Skills:** Airflow, Amazon Web Services, Amazon S3, Big Data, Cloud Computing, Information Engineering, Data Infrastructure, Data Systems, Data Warehousing, Distributed Data Store, Apache Hadoop, Python (Programming Language), Machine Learning, Query Optimization, SQL Databases, Unstructured Data, Apache Spark, Information Technology, Apache Flink, Apache Kafka, Presto - **Published:** June 28, 2026 - **Apply:** https://www.dice.com/job-detail/affded64-6152-46fb-9e3b-cee03c16a477 ## About the Role * 5+ years of data engineering experience, with at least 3 years owning production pipelines in a distributed data environment. * Strong hands-on experience with Spark and Kafka at scale - you've debugged a production incident in both, not just run tutorials. * Experience with big-data technologies including Hadoop, Presto/Trino, and Flink. * Proficiency in Python and SQL; Scala a plus. * Experience working in a cloud environment, preferably AWS (S3, EMR, Glue, Redshift). * Track record of improving efficiency, scalability, and stability of data systems - with measurable results. * BS or MS in Computer Science or equivalent. * Strong communication skills; comfortable driving data infrastructure decisions across application and platform teams. ## Description NewsBreak reaches tens of millions of users every day. Every feed ranking decision, every recommendation, every A/B test, and every ML model we ship runs on the data infrastructure you'd own. We're looking for a senior engineer to build and scale the batch and streaming pipelines at the core of that platform - someone who takes end-to-end ownership, drives reliability, and can translate ambiguous business needs into production-grade data systems. Responsibilities * Own the streaming data backbone - design and operate high-throughput Kafka pipelines carrying user events (clicks, reads, impressions, shares) from mobile clients through to downstream consumers: the data warehouse and real-time analytics. * Build and maintain batch pipelines at scale - author and optimize Spark jobs processing billions of rows daily: content ingestion, engagement aggregation, and user-level feature computation. Own pipeline reliability, incremental backfill strategies, and cost per TB processed. * Drive pipeline observability - instrument data quality checks, freshness monitors, and anomaly alerts so issues are caught before they reach dashboards. Define SLOs for data freshness and completeness, and own them. * Model structured and unstructured data - design data models that serve analytics and product use cases. Act as the data POC across teams and be the person who pushes back when a schema decision will hurt us six months later. * Partner with data scientists and platform engineers - understand data needs across teams and drive key data infrastructure decisions. Unblock downstream teams without becoming a ticket queue. * Improve efficiency - reduce compute costs and job latency through query optimization, smarter partitioning, better resource scheduling, and tiered storage strategies. * Raise the platform bar - contribute to shared infrastructure: pipeline framework standards, orchestration patterns (Airflow), reusable Spark libraries, and data catalog hygiene. Mentor junior engineers on the team ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Why and when should we consider Stream Processing frameworks in our solutions](https://www.wearedevelopers.com/videos/1085-why-and-when-should-we-consider-stream-processing-frameworks-in-our-solutions) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) ## Related Articles - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [A Guide to Green Tech and Green IT Careers](https://www.wearedevelopers.com/magazine/374-a-guide-to-green-tech-and-green-it-careers)