Sr Data Software Engineer

Info Dinamica Inc
East Coast of the United States, United States
5 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Amazon S3 Apache HTTP Server Big Data Data Governance Data Infrastructure Data Integrity Data Security Database Queries Distributed Computing Environment Python (Programming Language) Meta-Data Management
+8 more
Scripting Apache Spark Build Management Data Lakes AWS Glue Data Programming AWS Data Analytics Data Pipelines

Job description

Note: The strongest candidates are likely to come from organizations with mature AWS data lake ecosystems, such as large financial services, healthcare, retail, or cloud-native technology companies that have recently adopted Apache Iceberg for their data platform modernization efforts. Must-have skills: Hands-on data engineer with strong Python skills and experience building large-scale Apache Spark pipelines on AWS. Deep platform fluency across the modern data lake stack,S3, Glue Data Catalog, Athena, EMR, and Redshift, with a focus on reliable ingestion, transformation, and validation at scale. Project: The client has a central data platform built on AWS Glue Data Catalog. They’ve developed a fine-grained data access control (FGAC) system that is now being adopted across various data-producing teams. As part of this adoption, each application team must review their data pipelines to identify queries affected by the new access controls and modify the queries they own accordingly. Delta-to-Iceberg Migration

  • Design and build a migration feature/framework to convert tables from Delta Lake format to Apache Iceberg at scale.
  • Leverage deep understanding of Iceberg internals (metadata layers, manifests, snapshots, partitioning, schema evolution) to ensure correct and efficient migration
  • Validate data integrity, schema fidelity, and partitioning strategy parity between source Delta tables and migrated Iceberg tables
  • Optimize migration processes for performance and reliability across large volumes of tables and data
  • Identify and resolve edge cases, compatibility issues, or performance bottlenecks during migration at scale

FGAC (Fine-Grained Access Control)

  • Ensure Row-Level Security (RLS) and Column-Level Security (CLS) policies are correctly preserved or re-implemented on migrated Iceberg tables
  • Integrate migration workflows with AWS Lake Formation to maintain governance and access control continuity
  • Collaborate with security/governance teams to validate that FGAC policies enforce correctly post-migration
  • Support GDC (Global Data Catalog / Governed Data Catalog) integration to ensure migrated tables remain properly cataloged and governed
  • Build validation/reconciliation tooling to confirm access control parity between pre- and post-migration states

General Responsibilities

  • Document migration architecture, access control mapping logic, and operational runbooks
  • Collaborate with data platform, governance, and security teams throughout the migration lifecycle
  • Monitor and troubleshoot production issues related to migrated tables and access enforcement
  • Contribute to best practices and reusable tooling for future Delta-to-Iceberg migration phases

Requirements

  • Strong, hands-on knowledge of Apache Iceberg internals (table format, metadata management, snapshots, partitioning, schema evolution)
  • Proven experience performing Delta-to-Iceberg migrations at scale
  • Experience with AWS Lake Formation for access control and data governance
  • Practical experience implementing or preserving Row-Level Security (RLS) and Column-Level Security (CLS)
  • Experience with GDC (Global/Governed Data Catalog) or equivalent data cataloging/governance systems
  • Solid understanding of distributed data processing and large-scale data migration challenges
  • Strong SQL skills and experience with table format internals (Delta Lake, Iceberg, or similar), * Experience with Spark for large-scale data processing and migration tooling
  • Familiarity with AWS data platform services (S3, Glue, Athena)
  • Experience with data governance frameworks and compliance-driven access control requirements
  • Prior experience building reusable migration frameworks or tooling for table format conversions
  • Scripting/automation experience (Python, Scala) for migration validation and tooling

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

55 sec

Validating data processing architectures via containerized events

Modood Alvi · World Congress 2025

3:43 min

The enduring legacy of the amazon S3 storage API

Chris Heilmann +3 · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:10 min

Why organizations combine big data and machine learning

Ayon Roy · LIVE

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

Videos

See all

Related articles

See all