Iceberg DBA / Big Data Administrator / Lakehouse Operations Engineer

Next Level Business Services Inc
United States
24 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
1 year minimum
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Apache HTTP Server Microsoft Azure Big Data Cloudera Impala Data Architecture Data Validation Information Engineering Data Governance Data Infrastructure Extract Transform Load (ETL) Data Security
+18 more
Apache Hadoop Apache Hive Information Lifecycle Management Python (Programming Language) Metadata Meta-Data Management Performance Tuning Role-Based Access Control DataOps Cloudera Teradata SQL Parquet Scripting Apache Spark Software Troubleshooting Data Lakes Apache Nifi Data Pipelines

Job description

We are seeking a Senior Iceberg DBA / Lakehouse Operations Engineer to support the reliability, performance, availability, and operational integrity of an enterprise-scale Apache Iceberg Lakehouse environment., * Own day-to-day operational administration of Apache Iceberg tables

  • Maintain reliability, availability, and consistency of enterprise Lakehouse datasets
  • Perform Iceberg table maintenance and optimization
  • Manage:
  • Compaction
  • Small-file mitigation
  • Snapshot expiration
  • Metadata cleanup
  • Orphan-file cleanup
  • Partition evolution
  • File-size optimization
  • Troubleshoot Iceberg table and metadata issues across Spark, Hive and Impala
  • Monitor production workloads and resolve performance issues
  • Troubleshoot query failures, inefficient scans and execution-plan issues
  • Support large-scale Lakehouse environments at multi-TB/PB scale
  • Perform data validation and reconciliation between source and Iceberg datasets
  • Support Hive/Teradata modernization to Iceberg
  • Assist with schema and data-type alignment during migration
  • Provide L2/L3 production support
  • Participate in on-call rotation
  • Handle P1/P2 incidents and meet defined SLAs
  • Conduct RCA and implement preventive actions
  • Collaborate with Data Engineering, Platform, Application and Data Governance teams
  • Support Ranger policies, RBAC and secure data access
  • Maintain data lifecycle, retention and archival policies

Requirements

Candidates coming from a Hadoop Administrator, Cloudera Administrator, Big Data DBA, Data Platform Administrator, or Lakehouse Operations background who have hands-on Apache Iceberg administration and table operations are strongly preferred.

Mid-level candidates with solid hands-on administration and production support experience will be considered., The ideal candidate will have 4 6 years of Big Data / Data Operations / DBA experience, including 4+ years with the Cloudera ecosystem (CDP) and at least 1+ year of hands-on Apache Iceberg experience., * 4 6 years of experience in Big Data Administration, Data Operations, DBA, Hadoop Administration, or Data Platform Operations

  • 1+ year hands-on Apache Iceberg experience
  • 4+ years of experience with Cloudera / CDP ecosystem
  • Strong experience with Iceberg table administration and maintenance
  • Hands-on experience with:
  • Iceberg table operations
  • Table maintenance
  • Compaction
  • Small-file management
  • Snapshot expiration
  • Metadata management
  • Orphan file cleanup / vacuum
  • Partition management and optimization
  • Strong experience with Spark SQL, Hive and/or Impala
  • Production support and L2/L3 incident management
  • Strong troubleshooting and root-cause-analysis skills
  • Experience supporting large-scale data platforms
  • Knowledge of TB/PB-scale data environments
  • Understanding of data lake / Lakehouse architecture
  • Strong knowledge of:
  • Partitioning strategies
  • Parquet / ORC
  • Distributed query processing
  • Table-level performance optimization
  • Data lifecycle management
  • Experience with monitoring, alerting, troubleshooting and operational support
  • Experience with data validation, reconciliation and data consistency
  • Knowledge of Ranger, RBAC and data access controls
  • Strong scripting skills using Python and/or Shell

Preferred Skills

  • Cloudera CDP
  • CDE / CDW
  • Hadoop Administration
  • Hive Administration
  • Apache Iceberg
  • Trino
  • NiFi
  • AWS or Azure
  • Hive-to-Iceberg migration
  • Teradata-to-Iceberg migration
  • Experience supporting Bronze / Silver / Gold / Medallion architecture
  • Experience with Iceberg across multiple query engines, * Big Data Development
  • ETL Development
  • Spark Development
  • Python Data Engineering
  • Data Pipeline Development

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:02 min

Choosing the right data format and catalog engine

Matthias Niehoff Matthias Niehoff · World Congress 2026 Europe

2:00 min

Separating dataset creation from low-level software implementation steps

Jan Zawadzki · World Congress 2022

2:50 min

How Parquet metadata enables efficient data reading

Matthias Niehoff Matthias Niehoff · World Congress 2026 Europe

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:03 min

Introduction to open table formats built on Parquet

Matthias Niehoff Matthias Niehoff · World Congress 2026 Europe

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

Videos

See all

Related articles

See all