Lead Data Engineer

BDIPLUS INC.
New York, NY, United States
about 1 month ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$120,000.0 - $140,000.0
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) IBM AIX Amazon Web Services Amazon S3 Apache HTTP Server IBM System I COBOL (Programming Language) Computer Programming Information Engineering Data Governance Data Infrastructure Extract Transform Load (ETL)
+20 more
JSON Python (Programming Language) Node.Js Open Database Connectivity Query Optimization Standard Sql Virtual Storage Access Methods YAML Openapi Apache Spark Caching Apigee Fastapi Debezium Data Analytics Data Lakehouse Api Design Api Gateway Data Pipelines User Identification

Job description

  • Build governed virtual data models in Denodo, including cross-system joins, canonical schemas, ABAC policies, and column-level masking
  • Configure and optimize semantic layer connectivity (JDBC/ODBC, connection pooling, failover)
  • Implement data quality rules (completeness, validity, uniqueness, anomaly detection, scoring)
  • Configure data catalog, lineage, and marketplace publishing
  • Build physical data products: *

  • Source extraction (mainframe, CDC, COBOL copybooks)
  • ETL/ELT pipelines using Informatica IDMC or dbt
  • Data Lakehouse storage using Apache Iceberg on AWS S3
  • Define and implement machine-readable data contracts (YAML/JSON), including schema guarantees, SLAs, and dependency tracking
  • Develop APIs, SLA monitoring dashboards, reconciliation processes, and production-grade delivery patterns
  • Partner with client teams through demos, rotations, and structured knowledge transfer
  • Leverage AI-assisted development (Claude Code) to accelerate delivery and improve productivity

Requirements

  • 5+ years of data engineering experience in enterprise-scale environments
  • Strong expertise with Denodo (VQL development, query optimization, caching, governance)
  • Experience with Apache Iceberg on AWS (S3, Glue, Athena), including schema evolution and partitioning
  • Hands-on experience with Informatica
  • Experience implementing data quality frameworks
  • Strong SQL and programming skills (Python, Spark; dbt preferred)
  • Experience with CDC technologies (e.g., Debezium, Informatica CDC, or equivalent)
  • Familiarity with API development (FastAPI, Node.js) and OpenAPI standards
  • Experience working with AI-assisted development tools (Claude Code preferred; training available)

Preferred Qualifications

  • Experience with mainframe data environments (COBOL copybooks, VSAM, AIX/AS400 extraction)
  • Financial services or insurance domain experience
  • Familiarity with multi-tier data certification frameworks (e.g., Bronze/Silver/Gold or Foundation/Production/Enterprise)
  • Experience with API gateways and AI integration patterns (e.g., Kong Gateway, APIGEE, MCP)
  • Experience designing and enforcing formal data contracts and dependency management
  • Consulting or client-facing experience with structured knowledge transfer
  • AWS certifications (Solutions Architect, Data Analytics) are a plus

About the company

BDIPlus is seeking a Senior Data Engineer to support a Fortune 100 financial services client’s real-time intelligence initiatives., BDIPlus is a data and AI consulting firm with deep experience delivering for Fortune 500 clients, including American Express and Morgan Stanley. We specialize in enterprise data products, identity resolution, semantic layers, and event-driven architecture.

Our engineers are AI-native, leveraging tools like Claude Code from day one to deliver faster, higher-quality solutions than traditional consulting teams.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:47 min

Exploring JSON, CBOR, and JOSE for data serialization

Aaron Russell · LIVE

1:35 min

Centralizing configuration logic with native YAML block references

Matthieu Vincent Matthieu Vincent · Europe 2026 Virtual

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:03 min

Distinguishing type definition constructs from data validation routines

Clemens Vasters Clemens Vasters · World Congress 2025

2:00 min

Introduction to YAML syntax and basic formatting

Chris Ayers · LIVE

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

Videos

See all

Related articles

See all