Data Pipeline & Ingestion Engineer (Senior / Mid)

FANISKO LLC
United States
about 2 months ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Data Deduplication Extract Transform Load (ETL) Data Mapping Relational Databases Protocol Buffers JSON Python (Programming Language) Operational Databases Reference Data Standard Sql YAML
+6 more
Git Avro Apache Kafka Crosswalk Data Pipelines User Identification

Job description

You will build and operate the data backbone of ODL: bulk and streaming ingestion from legacy source systems, medallion-layered storage (Bronze/Silver/Gold), identity resolution and golden-record consolidation, source-to-canonical mapping and crosswalks, and the data-quality and reconciliation gates that prove data is complete and correct before it is published. This is the volume engine of the program - every new client onboarded flows through the pipelines you build., * Build batch-seed and event-tail ingestion per source system, including seed tail watermark hand-off,idempotent upserts, and dedup ledgers

  • Build and operate medallion layers with reprocess-from-Bronze, pipeline orchestration (checkpoints,retry/backoff, DLQ), and full observability
  • Build data-quality gates (quarantine / pass-with-flag), quality scoring, and a reconciliation engine covering count, record, and financial reconciliation - financial is zero-tolerance
  • Build identity matching combining deterministic rules with probabilistic scoring and confidence bands;deliver deduplication, golden-record materialization, and survivorship rules, calibrating match thresholds with labelled data
  • Author and maintain source canonical structural mappings and value crosswalks (e.g., collapsing 1,800+ raw employment-status values to ~20 standard ones) as governed, versioned configuration
  • Enforce data contracts at the boundary: schema registry, fail-fast validation, and semver-compatible schema evolution, * Probabilistic record linkage at depth - blocking/candidate generation, scoring models, threshold calibration (expected at senior level)
  • Schema registry experience (Avro/Protobuf)
  • Extracting from mainframe or older RDBMS sources with limited CDC support
  • Financial reconciliation in finance-adjacent domains
  • Benefits administration or healthcare domain knowledge

Requirements

  • 5+ years building production data pipelines at scale
  • Kafka depth: consumers/producers, replay, DLQ, exactly-once / idempotent processing patterns
  • Strong SQL and solid ETL fundamentals
  • Java and/or Python in production
  • Medallion / lakehouse layering, CDC, watermark/checkpoint patterns, and batch-stream hand-off
  • Data-quality frameworks: validation rules, quarantine and re-entry, quality scoring, reconciliation
  • Entity resolution / MDM exposure: record matching, dedup, survivorship - via commercial tools(Informatica MDM, Reltio) or custom builds
  • Data mapping and crosswalk discipline: profiling messy datasets, authoring governed reference data,config-as-code (YAML/JSON, Git)

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

47 sec

Building modern data pipelines for legacy exports

Dr. Alexander Wachtel Dr. Alexander Wachtel +1 · World Congress 2025

3:02 min

Audience Q&A on data formats and engine tradeoffs

Matthias Niehoff Matthias Niehoff · World Congress 2026 Europe

3:47 min

Exploring JSON, CBOR, and JOSE for data serialization

Aaron Russell · LIVE

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

1:52 min

Customizing block storage tiers and formats

Ricardo Sueiras Sueiras · LIVE

2:03 min

Distinguishing type definition constructs from data validation routines

Clemens Vasters Clemens Vasters · World Congress 2025

Videos

See all

Related articles

See all