Data Engineer -Iceberg

American IT Systems
Atlanta, GA, United States
12 days ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Compensation
$144,000.0 - $210,000.0
Working hours
Regular working hours

Tech stack

Airflow Apache HTTP Server Big Data Computer Programming Data Validation Information Engineering Extract Transform Load (ETL) Data Transformation Data Security Data Systems Distributed Computing Environment Python (Programming Language)
+15 more
Performance Tuning Standard Sql SQL Stored Procedures SQL Databases SQL Server Integration Services Data Logging Data Processing Freeform SQL Data Ingestion Apache Spark Software Troubleshooting Data Lakes Pyspark Data Pipelines Databricks

Job description

We are seeking an experienced Data Engineer with strong hands-on experience in SQL, Python, SSIS, Airflow, Apache Spark, Apache Iceberg, and Databricks. The ideal candidate will be responsible for designing, developing, and optimizing scalable data pipelines and data processing solutions in a cloud-based environment., * Design, develop, and maintain scalable ETL/ELT data pipelines using Python, SQL, SSIS, Airflow, and Spark.

  • Develop complex SQL queries, stored procedures, data transformations, and performance optimization solutions.
  • Build and maintain data pipelines and workflows using Apache Airflow, including scheduling, dependencies, monitoring, retries, and failure handling.
  • Develop and optimize Apache Spark/PySpark applications for large-scale data processing.
  • Work with Databricks to develop, deploy, and optimize data engineering workloads.
  • Work with Apache Iceberg tables and modern data lake/lakehouse architectures.
  • Migrate and modernize legacy SSIS ETL workflows into cloud-based and Spark/Databricks-based data pipelines.
  • Implement data ingestion, transformation, validation, and integration processes from multiple sources.
  • Optimize Spark jobs, SQL queries, data partitioning, and storage strategies for performance and scalability.
  • Implement data quality checks, error handling, logging, monitoring, and pipeline recovery mechanisms.
  • Collaborate with data architects, developers, analysts, and business teams to understand requirements and deliver reliable data solutions.
  • Follow best practices for data security, governance, performance, and operational reliability.

Requirements

  • Strong hands-on experience with SQL
  • Strong programming experience with Python
  • Hands-on experience with Apache Spark / PySpark
  • Experience with Databricks
  • Strong experience with Apache Airflow
  • Hands-on experience with SSIS
  • Experience with Apache Iceberg
  • Strong understanding of ETL/ELT and data pipeline development
  • Experience working with large-scale datasets and distributed data processing
  • Strong troubleshooting and performance-tuning skills
  • Good understanding of modern data lake/lakehouse architectures

About the company

Cox Automotive

  • Atlanta, GA
  • $163,400 per year Cox Automotive is deploying enterprise AI capabilities on AWS Quick across the enterprise, helping teams work more effectively in their day-to-day operations. We are seeking a Pr…, Cargill

  • Atlanta, GA
  • $144,000-210,000 per year Cargill’s size and scale allows us to make a positive impact in the world. Our purpose is to nourish the world in a safe, responsible and sustainable way. Cargill is a family comp…, Cargill is committed to providing food and agricultural solutions to nourish the world in a safe, responsible, and sustainable way. Sitting at the heart of the supply chain, we par…

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

2:10 min

Why organizations combine big data and machine learning

Ayon Roy · LIVE

2:34 min

Capabilities of the Apache Spark processing engine

Ayon Roy · LIVE

Videos

See all

Related articles

See all