> Markdown version of [/jobs/ext/1281377-data-pipeline-ingestion-engineer-senior-mid](https://www.wearedevelopers.com/jobs/ext/1281377-data-pipeline-ingestion-engineer-senior-mid). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Pipeline & Ingestion Engineer (Senior / Mid) - **Company:** FANISKO LLC - **Location:** United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Data Deduplication, Extract Transform Load (ETL), Data Mapping, Relational Databases, Protocol Buffers, JSON, Python (Programming Language), Operational Databases, Reference Data, Standard Sql, YAML, Git, Avro, Apache Kafka, Crosswalk, Data Pipelines, User Identification - **Published:** July 15, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=7a5ce8106aad85d6 ## About the Role * 5+ years building production data pipelines at scale * Kafka depth: consumers/producers, replay, DLQ, exactly-once / idempotent processing patterns * Strong SQL and solid ETL fundamentals * Java and/or Python in production * Medallion / lakehouse layering, CDC, watermark/checkpoint patterns, and batch-stream hand-off * Data-quality frameworks: validation rules, quarantine and re-entry, quality scoring, reconciliation * Entity resolution / MDM exposure: record matching, dedup, survivorship - via commercial tools(Informatica MDM, Reltio) or custom builds * Data mapping and crosswalk discipline: profiling messy datasets, authoring governed reference data,config-as-code (YAML/JSON, Git) ## Description You will build and operate the data backbone of ODL: bulk and streaming ingestion from legacy source systems, medallion-layered storage (Bronze/Silver/Gold), identity resolution and golden-record consolidation, source-to-canonical mapping and crosswalks, and the data-quality and reconciliation gates that prove data is complete and correct before it is published. This is the volume engine of the program - every new client onboarded flows through the pipelines you build., * Build batch-seed and event-tail ingestion per source system, including seed tail watermark hand-off,idempotent upserts, and dedup ledgers * Build and operate medallion layers with reprocess-from-Bronze, pipeline orchestration (checkpoints,retry/backoff, DLQ), and full observability * Build data-quality gates (quarantine / pass-with-flag), quality scoring, and a reconciliation engine covering count, record, and financial reconciliation - financial is zero-tolerance * Build identity matching combining deterministic rules with probabilistic scoring and confidence bands;deliver deduplication, golden-record materialization, and survivorship rules, calibrating match thresholds with labelled data * Author and maintain source canonical structural mappings and value crosswalks (e.g., collapsing 1,800+ raw employment-status values to ~20 standard ones) as governed, versioned configuration * Enforce data contracts at the boundary: schema registry, fail-fast validation, and semver-compatible schema evolution, * Probabilistic record linkage at depth - blocking/candidate generation, scoring models, threshold calibration (expected at senior level) * Schema registry experience (Avro/Protobuf) * Extracting from mainframe or older RDBMS sources with limited CDC support * Financial reconciliation in finance-adjacent domains * Benefits administration or healthcare domain knowledge ## Related Videos - [Implementing continuous delivery in a data processing pipeline](https://www.wearedevelopers.com/videos/73-implementing-continuous-delivery-in-a-data-processing-pipeline) - [From event streaming to event sourcing 101](https://www.wearedevelopers.com/videos/91-from-event-streaming-to-event-sourcing-101) - [Tips and Tricks for Working with JSON](https://www.wearedevelopers.com/videos/1229-tips-and-tricks-for-working-with-json) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Introducing JSON Structure](https://www.wearedevelopers.com/videos/100219-introducing-json-structure) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Dev Digest 159: AI Pipelines, 10x Faster TypeScript, How to Interview](https://www.wearedevelopers.com/magazine/563-dev-digest-159-ai-pipelines-10x-faster-typescript-how-to-interview) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering)