Data Quality Engineer

Select Minds LLC
Dallas, TX, United States
2 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Amazon S3 Continuous Integration Data as a Services Data Validation Information Engineering Data Integrity Extract Transform Load (ETL) Serialization Data Systems Database Queries Database Testing
+23 more
Software Debugging Distributed Data Store Github Protocol Buffers JSON Python (Programming Language) Prometheus SQL Databases Data Streaming Grafana Concurrency Apache Spark Pyspark Data Lineage Low Latency Avro Apache Kafka Data Management Cloudwatch Data Pipelines SDET Jenkins Databricks

Job description

We are looking for a Data Quality Engineer to own validation across batch and streaming data pipelines. This role focuses on ensuring data correctness, reliability, and performance across platforms built on Databricks, Kafka, AWS, SQL, and Python. This is a hands-on role focused on building scalable data validation frameworks and ensuring production-grade data systems. Key Responsibilities End-to-End Data Validation

  • Validate data pipelines for accuracy, completeness, consistency, and timeliness
  • Build SQL-based validations for business rules and transformations
  • Implement reconciliation between source and downstream systems
  • Ensure data lineage and traceability ETL / ELT & Spark Testing

  • Test pipelines built on AWS (Glue, Lambda, EMR, Step Functions)
  • Validate transformations using SQL and Python
  • Test ingestion, transformation, aggregation, and serving layers
  • Handle backfills, reprocessing, and historical data loads
  • Validate Spark pipelines (PySpark/Scala) on Databricks Streaming (Kafka)

  • Validate data integrity, ordering, and delivery guarantees
  • Test producer and consumer logic and serialization formats (Avro, JSON, Protobuf)
  • Validate topics, partitions, offsets, retention, and schema evolution
  • Simulate late events, duplicates, and failure scenarios Automation & Frameworks

  • Build Python-based data testing frameworks
  • Develop reusable validation utilities and synthetic datasets
  • Integrate data tests into CI/CD pipelines
  • Enable automated alerts for data quality issues Performance & Reliability

  • Validate throughput, latency, and concurrency at scale
  • Test retry logic, idempotency, and recovery mechanisms
  • Perform regression, soak, and failover testing Observability

  • Validate logs, metrics, and alerts using tools such as CloudWatch, Prometheus, and Grafana
  • Define and monitor data SLAs and SLOs
  • Support incident response, root cause analysis, and postmortems

Requirements

  • 7+ years of total experience in QA, SDET, or Data Quality Engineering
  • Minimum 4-6 years of hands-on experience working with data platforms, data pipelines, or data engineering ecosystems
  • 3+ years of hands-on experience with Databricks and Apache Spark
  • Strong SQL skills for data validation, reconciliation, and complex analysis
  • Proficiency in Python for automation and data validation
  • Experience testing ETL/ELT pipelines (batch and streaming)
  • Hands-on experience with Kafka or similar streaming platforms
  • Strong understanding of AWS data services (S3, Glue, Lambda, Redshift, Athena)
  • Experience working with large-scale distributed data systems
  • Strong debugging, analytical, and problem-solving skills Nice to Have

  • Experience with data quality or observability tools such as Great Expectations or Monte Carlo
  • Knowledge of schema registry and data contracts
  • Experience with CI/CD tools such as GitHub Actions or Jenkins Flexible work from home options available.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.wayup.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

3:02 min

Audience Q&A on data formats and engine tradeoffs

Matthias Niehoff Matthias Niehoff · WWC Europe 2026

3:47 min

Exploring JSON, CBOR, and JOSE for data serialization

Aaron Russell · LIVE

1:52 min

Customizing block storage tiers and formats

Ricardo Sueiras Sueiras · LIVE

2:03 min

Distinguishing type definition constructs from data validation routines

Clemens Vasters Clemens Vasters · WWC 2025

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · WWC Europe 2026

Videos

See all

Related articles

See all