Data Engineer

Zeta Global
United States
4 days ago
Apply on job-boards.greenhouse.io
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Compensation
$139,000.0 - $219,000.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Big Data Command-Line Interface Data Architecture Data Validation Information Engineering Data Governance Data Infrastructure Data Integration Extract Transform Load (ETL) Software Debugging
+14 more
Distributed Computing Environment Python (Programming Language) Operational Databases Software Tools Standard Sql Software Engineering Data Ingestion Apache Spark Gitlab Git Data Management Api Design Data Pipelines Databricks

Job description

Lead design and delivery of scalable production data pipelines and data products (Spark/Python/SQL) on Databricks; integrate diverse federal health systems, define reusable ingestion patterns, ensure data quality/governance, mentor engineers, and troubleshoot/optimize production workflows., In this role, you’ll serve as a senior technical contributor, delivering scalable, production-ready data capabilities while establishing reusable patterns for data ingestion, integration, and delivery. You’ll solve complex data challenges, guide engineering teams, and help shape the technical foundation of a growing data platform ecosystem supporting federal health missions., * Design, develop, and maintain scalable, production-ready data pipelines and data products using Spark (Python/SQL) in a Databricks environment

  • Lead the integration and transformation of complex data from diverse DoW and federal health systems and sources into reliable, reusable data products
  • Design scalable approaches for data ingestion, integration, and exchange, including API-based integrations and services
  • Establish and promote reusable data engineering patterns, standards, and best practices that improve consistency, scalability, and maintainability across data products
  • Provide technical guidance on data architecture, pipeline design, data modeling, integration approaches, and engineering practices
  • Monitor, troubleshoot, and optimize production workflows and data pipelines, identifying performance, reliability, and scalability improvements
  • Define and implement data validation, quality, and governance practices that improve the reliability and usability of data products
  • Troubleshoot complex technical and data integration challenges, identify root causes, and drive sustainable solutions
  • Collaborate with engineers, architects, analysts, and customer stakeholders to translate complex data needs into scalable technical solutions
  • Provide technical guidance and mentorship to other engineers, helping teams navigate complex or unfamiliar technical challenges
  • Proactively identify opportunities to improve engineering tools, processes, and patterns and help drive their adoption across the team
  • Take ownership of complex technical areas and help maintain engineering quality, consistency, and cohesion as the platform and portfolio of data products grow, Build and maintain ETL/ELT pipelines to move product data into analytics-ready stores (Postgres, data lake, warehouse). Design and optimize data models, ensure data quality and documentation, support ad-hoc research requests, and collaborate with research and engineering teams to enable reproducible analytics.

Requirements

Citizenship & Clearance Requirement: per client requirements, candidates must be U.S. Citizens with an active DoW Secret (or higher) clearance Education Requirement: Bachelor’s Degree in Computer Science, Engineering, Data Science, or a related technical field (preferred) 540 Internal Thrive Level: Senior Data Engineer, * 10+ years of data engineering, software engineering, or related technical experience

  • Extensive hands-on experience designing, building, and operating production data pipelines and data products
  • Advanced proficiency with Python and SQL
  • Strong experience with Apache Spark and distributed data processing
  • Experience working with Databricks or similar modern data platforms
  • Experience designing and maintaining ETL/ELT processes for complex, large-scale datasets
  • Experience integrating data across disparate systems and consuming or developing API-based data integrations
  • Strong understanding of data modeling, data architecture, data quality, and data governance principles
  • Experience troubleshooting and optimizing complex production data pipelines for performance, reliability, and scalability
  • Experience working with Git-based development workflows and modern software engineering practices
  • Experience working in terminal / command-line environments
  • Demonstrated experience providing technical guidance, mentoring engineers, and influencing engineering practices
  • Strong client and stakeholder communication skills, with the ability to translate technical concepts and recommendations for both technical and non-technical audiences
  • Ability to independently navigate ambiguity, identify technical risks, and drive complex engineering challenges toward resolution
  • Experience spotting security, privacy and compliance issues and working with security/compliance/legal stakeholders

NICE TO HAVE SKILLS & EXPERIENCE

  • AWS cloud experience
  • Experience working with very large datasets, including datasets with billions of records
  • Experience with Palantir Foundry
  • GitLab experience
  • Experience working with Advana or similar DoW data environments
  • Experience working with federal health, financial, or other regulated and sensitive data
  • Experience using AI/ML to accelerate work, including automating routine tasks, accelerating development and debugging

Benefits & conditions

  • Flexible PTO + all Federal holidays off
  • Health, dental and vision insurance plans
  • Flexible Spending Account (FSA)
  • 401k with employer match
  • Company-sponsored life insurance, short- and long-term disability
  • Professional development (training, certifications, conferences)
  • Paid cloud developer accounts
  • Referral bonuses
  • HQ office perks (parking / metro reimbursement, nitro coffee & lunches)
  • Annual social events (540 Week, hackathon, charity golf tournament, etc.)
  • Access to 540’s Washington Capitals & Nationals tickets, 4 Days Ago In-Office or Remote 175K-220K Annually Senior level 175K-220K Annually Senior level Artificial Intelligence * Cloud * Software * Infrastructure as a Service (IaaS) Design and implement scalable data pipelines and a SOC2-compliant data warehouse. Support analytics and ML teams, enable real-time and batch ETL, mentor data engineers, and collaborate cross-functionally to drive data-driven decisions. Top Skills: DagsterDatabricksDbtLlmsPlanetscaleRedshiftSnowflakeSoc2Tinybird Zeta Global, United States Easy Apply 140K-160K Annually Senior level 140K-160K Annually Senior level AdTech * Artificial Intelligence * Marketing Tech * Software * Analytics Build, deploy, and operate production-grade data pipelines and data products for healthcare audiences. Design transformations, data models, and governed views using Python, SQL, Airflow, S3, Snowflake, and EMR. Implement data-quality, monitoring, and privacy-by-design controls for PHI/PII. Partner with product, analytics, and platform teams to onboard sources, support audience discovery, segmentation, activation, measurement, and troubleshoot production issues. Top Skills: Amazon EmrAmazon S3Apache AirflowAthenaHivePythonSnowflakeSQL

What you need to know about the Colorado Tech Scene

With a business-friendly climate and research universities like CU Boulder and Colorado State, Colorado has made a name for itself as a startup ecosystem. The state boasts a skilled workforce and high quality of life thanks to its affordable housing, vibrant cultural scene and unparalleled opportunities for outdoor recreation. Colorado is also home to the National Renewable Energy Laboratory, helping cement its status as a hub for renewable energy innovation.

Key Facts About Colorado Tech

  • Number of Tech Workers: 260,000; 8.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Lockheed Martin, Century Link, Comcast, BAE Systems, Level 3
  • Key Industries: Software, artificial intelligence, aerospace, e-commerce, fintech, healthtech
  • Funding Landscape: $4.9 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Access Venture Partners, Ridgeline Ventures, Techstars, Blackhorn Ventures
  • Research Centers and Universities: Colorado School of Mines, University of Colorado Boulder, University of Denver, Colorado State University, Mesa Laboratory, Space Science Institute, National Center for Atmospheric Research, National Renewable Energy Laboratory, Gottlieb Institute

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on job-boards.greenhouse.io
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

6:14 min

Structuring CI/CD pipelines with integrated security and quality checks

Christoph Ruggenthaler · LIVE

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

Videos

See all

Related articles

See all