> Markdown version of [/jobs/ext/2104198-data-engineer](https://www.wearedevelopers.com/jobs/ext/2104198-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** Zeta Global - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $139,000.0 - $219,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Big Data, Command-Line Interface, Data Architecture, Data Validation, Information Engineering, Data Governance, Data Infrastructure, Data Integration, Extract Transform Load (ETL), Software Debugging, Distributed Computing Environment, Python (Programming Language), Operational Databases, Software Tools, Standard Sql, Software Engineering, Data Ingestion, Apache Spark, Gitlab, Git, Data Management, Api Design, Data Pipelines, Databricks - **Published:** August 18, 2026 - **Apply:** https://job-boards.greenhouse.io/540/jobs/7886221003 ## About the Role Citizenship & Clearance Requirement: per client requirements, candidates must be U.S. Citizens with an active DoW Secret (or higher) clearance Education Requirement: Bachelor's Degree in Computer Science, Engineering, Data Science, or a related technical field (preferred) 540 Internal Thrive Level: Senior Data Engineer, * 10+ years of data engineering, software engineering, or related technical experience * Extensive hands-on experience designing, building, and operating production data pipelines and data products * Advanced proficiency with Python and SQL * Strong experience with Apache Spark and distributed data processing * Experience working with Databricks or similar modern data platforms * Experience designing and maintaining ETL/ELT processes for complex, large-scale datasets * Experience integrating data across disparate systems and consuming or developing API-based data integrations * Strong understanding of data modeling, data architecture, data quality, and data governance principles * Experience troubleshooting and optimizing complex production data pipelines for performance, reliability, and scalability * Experience working with Git-based development workflows and modern software engineering practices * Experience working in terminal / command-line environments * Demonstrated experience providing technical guidance, mentoring engineers, and influencing engineering practices * Strong client and stakeholder communication skills, with the ability to translate technical concepts and recommendations for both technical and non-technical audiences * Ability to independently navigate ambiguity, identify technical risks, and drive complex engineering challenges toward resolution * Experience spotting security, privacy and compliance issues and working with security/compliance/legal stakeholders NICE TO HAVE SKILLS & EXPERIENCE * AWS cloud experience * Experience working with very large datasets, including datasets with billions of records * Experience with Palantir Foundry * GitLab experience * Experience working with Advana or similar DoW data environments * Experience working with federal health, financial, or other regulated and sensitive data * Experience using AI/ML to accelerate work, including automating routine tasks, accelerating development and debugging ## Description Lead design and delivery of scalable production data pipelines and data products (Spark/Python/SQL) on Databricks; integrate diverse federal health systems, define reusable ingestion patterns, ensure data quality/governance, mentor engineers, and troubleshoot/optimize production workflows., In this role, you'll serve as a senior technical contributor, delivering scalable, production-ready data capabilities while establishing reusable patterns for data ingestion, integration, and delivery. You'll solve complex data challenges, guide engineering teams, and help shape the technical foundation of a growing data platform ecosystem supporting federal health missions., * Design, develop, and maintain scalable, production-ready data pipelines and data products using Spark (Python/SQL) in a Databricks environment * Lead the integration and transformation of complex data from diverse DoW and federal health systems and sources into reliable, reusable data products * Design scalable approaches for data ingestion, integration, and exchange, including API-based integrations and services * Establish and promote reusable data engineering patterns, standards, and best practices that improve consistency, scalability, and maintainability across data products * Provide technical guidance on data architecture, pipeline design, data modeling, integration approaches, and engineering practices * Monitor, troubleshoot, and optimize production workflows and data pipelines, identifying performance, reliability, and scalability improvements * Define and implement data validation, quality, and governance practices that improve the reliability and usability of data products * Troubleshoot complex technical and data integration challenges, identify root causes, and drive sustainable solutions * Collaborate with engineers, architects, analysts, and customer stakeholders to translate complex data needs into scalable technical solutions * Provide technical guidance and mentorship to other engineers, helping teams navigate complex or unfamiliar technical challenges * Proactively identify opportunities to improve engineering tools, processes, and patterns and help drive their adoption across the team * Take ownership of complex technical areas and help maintain engineering quality, consistency, and cohesion as the platform and portfolio of data products grow, Build and maintain ETL/ELT pipelines to move product data into analytics-ready stores (Postgres, data lake, warehouse). Design and optimize data models, ensure data quality and documentation, support ad-hoc research requests, and collaborate with research and engineering teams to enable reproducible analytics. ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [WeAreDevelopers LIVE - Modern DevOps for IoT Devices and More](https://www.wearedevelopers.com/videos/1805-wearedevelopers-live-modern-devops-for-iot-devices-and-more) - [Data Governance in the Era of AI](https://www.wearedevelopers.com/videos/1622-data-governance-in-the-era-of-ai) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) ## Related Articles - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Data Analyst Salary in the UK](https://www.wearedevelopers.com/magazine/278-data-analyst-salary-in-the-uk)