Software Engineer - Data Acquisition Team
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+27 more
Job description
store large raw data sets from a wide range of sources. You will work across distributed processing, workflow orchestration, Architect resilient ETL/ELT workflows for batch, streaming, scheduled, and event-driven data processing.
- Develop production Java services and data processing applications for ingestion, orchestration, enrichment, deduplication, and
delivery.
- Build and improve systems using technologies such as Apache Airflow, Apache Beam, Spark, Google Dataflow, DataProc, Kafka,
and Pub/Sub.
- Define practical approaches for schema evolution, data contracts, data validation, backfills, replayability, and idempotent
processing.
-
Improve reliability, performance, scalability, and cost efficiency across data acquisition pipelines and services.
-
Implement observability, monitoring, and alerting for pipeline health, throughput, latency, failure rates, and data quality metrics.
-
Work with product, data science, platform, and data quality teams to translate business needs into production-ready systems.
-
Contribute to technical designs, implementation plans, and system modernization efforts across Data Acquisition.
Requirements
This is a senior individual-contributor engineering role focused on technical depth, production execution, and high-quality data
systems. The right candidate brings strong backend engineering fundamentals, significant pipeline experience, and the ability to
turn complex data acquisition requirements into scalable systems., 5+ years of professional software engineering experience with a strong focus on backend systems, data engineering, or distributed
processing.
-
Proven experience building and operating production data pipelines at scale.
-
Deep proficiency with Java and object-oriented design.
-
Hands-on expertise with data processing and orchestration technologies such as Apache Beam, Apache Airflow, Spark, Google
Dataflow, or DataProc.
-
Strong experience with streaming systems such as Apache Kafka, Google Pub/Sub, or similar technologies.
-
Strong understanding of batch processing, streaming processing, data modeling, schema evolution, and data quality
management.
ZoomInfo - Data Acquisition Page 1- Experience designing ETL/ELT workflows that process large volumes of structured and semi-structured data.
Backend and Distributed Systems
-
Experience designing high-throughput, fault-tolerant backend services and distributed systems.
-
Strong understanding of APIs, integration patterns, retries, backpressure, idempotency, and operational failure modes.
-
Ability to write clean, maintainable production code and evaluate tradeoffs in system design.
-
Experience with large-scale storage and query technologies such as BigQuery, Snowflake, Trino, or similar systems.
Cloud and Operations
-
Experience with at least one cloud provider, preferably Google Cloud Platform.
-
Hands-on experience with cloud services such as BigQuery, GCS, GKE, Dataflow, DataProc, and Pub/Sub.
-
Experience operating production services with monitoring, logging, alerting, SLIs, and incident investigation.
-
Ability to analyze and improve pipeline performance, reliability, resource utilization, and cost.
General
-
Bachelor’s degree in Computer Science, Software Engineering, or a related field.
-
Ability to translate business and data requirements into clear technical designs and implementation plans.
-
Pragmatic engineering judgment with the ability to balance correctness, delivery speed, maintainability, and operational risk.
Nice to Have
-
Experience with Kubernetes, especially GKE or EKS, for running distributed workloads.
-
Experience with Terraform or other infrastructure-as-code tools.
-
Experience with Snowflake, BigQuery, Starburst/Trino, or similar query engines.
-
Knowledge of data integration patterns involving CRM systems, email/calendar APIs, third-party feeds, change data capture, or
external data providers.
-
Experience in a B2B data company, data marketplace, or data-as-a-product environment.
-
Expert-level experience with Apache Spark or another distributed processing framework.
About the company
Data Acquisition is one of ZoomInfo’s biggest assets. The pipelines and services you build will directly shape the scale, ZoomInfo (NASDAQ: GTM) is the Go-To-Market Intelligence Platform that empowers businesses to grow faster with AI-ready insights, trusted data, and advanced automation. Its solutions provide more than 35,000 companies worldwide with a complete view of their customers, making every seller their best seller.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.dice.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Making Data Warehouses Fast: A Developer’s Story
Top Big Data Technologies That You Need to Know
Data Engineer Salary UK
Is Software Engineering Over-Saturated?