Iceberg DBA / Big Data Administrator / Lakehouse Operations Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+18 more
Job description
We are seeking a Senior Iceberg DBA / Lakehouse Operations Engineer to support the reliability, performance, availability, and operational integrity of an enterprise-scale Apache Iceberg Lakehouse environment., * Own day-to-day operational administration of Apache Iceberg tables
- Maintain reliability, availability, and consistency of enterprise Lakehouse datasets
- Perform Iceberg table maintenance and optimization
- Manage:
- Compaction
- Small-file mitigation
- Snapshot expiration
- Metadata cleanup
- Orphan-file cleanup
- Partition evolution
- File-size optimization
- Troubleshoot Iceberg table and metadata issues across Spark, Hive and Impala
- Monitor production workloads and resolve performance issues
- Troubleshoot query failures, inefficient scans and execution-plan issues
- Support large-scale Lakehouse environments at multi-TB/PB scale
- Perform data validation and reconciliation between source and Iceberg datasets
- Support Hive/Teradata modernization to Iceberg
- Assist with schema and data-type alignment during migration
- Provide L2/L3 production support
- Participate in on-call rotation
- Handle P1/P2 incidents and meet defined SLAs
- Conduct RCA and implement preventive actions
- Collaborate with Data Engineering, Platform, Application and Data Governance teams
- Support Ranger policies, RBAC and secure data access
- Maintain data lifecycle, retention and archival policies
Requirements
Candidates coming from a Hadoop Administrator, Cloudera Administrator, Big Data DBA, Data Platform Administrator, or Lakehouse Operations background who have hands-on Apache Iceberg administration and table operations are strongly preferred.
Mid-level candidates with solid hands-on administration and production support experience will be considered., The ideal candidate will have 4 6 years of Big Data / Data Operations / DBA experience, including 4+ years with the Cloudera ecosystem (CDP) and at least 1+ year of hands-on Apache Iceberg experience., * 4 6 years of experience in Big Data Administration, Data Operations, DBA, Hadoop Administration, or Data Platform Operations
- 1+ year hands-on Apache Iceberg experience
- 4+ years of experience with Cloudera / CDP ecosystem
- Strong experience with Iceberg table administration and maintenance
- Hands-on experience with:
- Iceberg table operations
- Table maintenance
- Compaction
- Small-file management
- Snapshot expiration
- Metadata management
- Orphan file cleanup / vacuum
- Partition management and optimization
- Strong experience with Spark SQL, Hive and/or Impala
- Production support and L2/L3 incident management
- Strong troubleshooting and root-cause-analysis skills
- Experience supporting large-scale data platforms
- Knowledge of TB/PB-scale data environments
- Understanding of data lake / Lakehouse architecture
- Strong knowledge of:
- Partitioning strategies
- Parquet / ORC
- Distributed query processing
- Table-level performance optimization
- Data lifecycle management
- Experience with monitoring, alerting, troubleshooting and operational support
- Experience with data validation, reconciliation and data consistency
- Knowledge of Ranger, RBAC and data access controls
- Strong scripting skills using Python and/or Shell
Preferred Skills
- Cloudera CDP
- CDE / CDW
- Hadoop Administration
- Hive Administration
- Apache Iceberg
- Trino
- NiFi
- AWS or Azure
- Hive-to-Iceberg migration
- Teradata-to-Iceberg migration
- Experience supporting Bronze / Silver / Gold / Medallion architecture
- Experience with Iceberg across multiple query engines, * Big Data Development
- ETL Development
- Spark Development
- Python Data Engineering
- Data Pipeline Development
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Making Data Warehouses Fast: A Developer’s Story
Highest Paying Tech Companies for Developers
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
Data Engineer Salary UK