> Markdown version of [/jobs/ext/734408-data-engineer-league-analytics-infrastructure](https://www.wearedevelopers.com/jobs/ext/734408-data-engineer-league-analytics-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer, League Analytics & Infrastructure - **Company:** Major League Baseball - **Location:** New York, NY, United States - **Experience:** Experienced - **Salary:** $115,000.0 - $140,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, Amazon Web Services, Data Analysis, Microsoft Azure, BigQuery, Code Review, Continuous Integration, Data Governance, Software Debugging, Data Flow Control, Github, Operational Databases, SQL Databases, Adobe FreeHand, Scripting, Google Cloud, Delivery Pipeline, Git, Infrastructure Automation Frameworks, Information Technology, Terraform - **Published:** June 29, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=5e63e2ea97ff196b ## About the Role Do you have experience in SQL?, * 2-4 years of production data engineering experience * Expert-level SQL - comfortable writing complex freehand queries (sub-queries, nested logic, window functions) and reading someone else's to spot issues * Strong Python for data processing, scripting, and automation * Hands-on dbt experience - you've built models across staging, intermediate, and mart layers, written tests, and shipped to production * Production Airflow experience - DAG authoring, dependency management, debugging failed runs * Deep familiarity with Google Cloud Platform (BigQuery, GCS, Pub/Sub) or equivalent depth in AWS/Azure with willingness to convert * Git-based development workflows - branches, PRs, code review as a daily practice * You communicate clearly with both engineers and non-engineers, take feedback well, and give it kindly * Execution mindset. You can own a project from requirements to deployment with minimal oversight. Nice-to-Have * A degree in Computer Science, Engineering, or a related field - or non-traditional background with equivalent practical experience * Experience with Terraform or other Infrastructure-as-Code tools * Experience with AI-assisted development or enterprise AI tooling (Gemini Enterprise, Vertex AI). We're early but ambitious - we see AI as a lever for engineering efficiency * A passion for baseball, or prior experience in sports, media, or entertainment * Ability to build creative solutions for unusual problems ## Description * Build production-grade pipelines using Airflow and dbt to orchestrate batch and streaming transformations across GCP, so that downstream analysts and engineers can trust the data they query without checking the wiring * Architect clean, layered data models (staging intermediate mart) that serve as the single source of truth for league analytics, applying dbt best practices for materialization, testing, and documentation * Operate the ingestion layer using Pub/Sub, GCS, Dataflow, and Knowledge Catalog DataPlex) to land both batch and streaming sources cleanly into the lakehouse * Implement observability and monitoring standards so that data quality issues surface before stakeholders notice them, not after * Manage code through GitHub-based CI/CD, contributing to the deployment workflows that keep our platform reliable and our changes safe * Adhere to data governance practices that keep proprietary baseball data secure and compliant ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 159: AI Pipelines, 10x Faster TypeScript, How to Interview](https://www.wearedevelopers.com/magazine/563-dev-digest-159-ai-pipelines-10x-faster-typescript-how-to-interview) - [Dev Digest 132 - Binging WADFlix?](https://www.wearedevelopers.com/magazine/473-dev-digest-132-binging-wadflix) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)