ETL Lead - DataStage CP4D | AWS Glue

IntraEdge, Inc.
Detroit, MI, United States
19 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Amazon S3 Data Analysis Big Data Cloud Computing Code Review Continuous Integration Data Governance Extract Transform Load (ETL) Data Transformation Data Warehousing
+21 more
IBM InfoSphere DataStage Python (Programming Language) Meta-Data Management Performance Tuning Scrum Methodology Standard Sql User-Centered Design Delivery Pipeline Snowflake Parallel Computation AWS Lambda Git Pyspark Semi-structured Data Data Lineage Collibra AWS Glue Data Analytics Software Version Control Serverless Computing Jenkins

Job description

We are seeking a highly experienced ETL Lead to design, lead, and optimize enterprise data integration workflows using IBM DataStage on Cloud Pak for Data (CP4D), AWS Glue & Lambda, and Snowflake. The ideal candidate will drive modern data transformation strategies for large-scale data ingestion pipelines supporting analytics, governance, and AI/ML workloads., * Lead the end-to-end design and implementation of ETL workflows for structured and semi-structured data from various sources (SFTP, DB2, Oracle, APIs).

  • Architect and maintain data ingestion and transformation pipelines using IBM DataStage (CP4D) and AWS Glue/Lambda functions.
  • Optimize data load performance and manage large data volumes with effective partitioning, incremental loads, and error handling.
  • Collaborate with cloud engineers and data modelers to ensure data is curated for consumption in Snowflake (Silver/Gold/Platinum layers).
  • Ensure data lineage, metadata management, and data quality in compliance with enterprise data governance standards.
  • Partner with governance teams to ensure integration with Collibra, or other metadata and privacy tools.
  • Support CI/CD integration, parameterization, and version control of ETL code via Git and DevOps pipelines.
  • Lead and mentor a team of onshore/offshore ETL developers; establish best practices and code review processes.
  • Troubleshoot production ETL issues and participate in on-call rotations as needed.

Requirements

  • 8+ years of enterprise data integration experience with at least 3+ years as a technical lead.
  • Strong expertise in IBM DataStage, including modern deployments on Cloud Pak for Data (CP4D).
  • Solid experience with AWS Glue (PySpark) and AWS Lambda for scalable, serverless data transformation.
  • Proven knowledge of Snowflake data warehousing, including external stages, streams, tasks, and performance tuning.
  • Experience with complex ETL orchestration involving parallel processing, dynamic parameterization, and error handling.
  • Strong SQL, Python (preferred for Glue), and performance tuning skills.
  • Understanding of data lakehouse architectures, S3, and ingestion patterns (batch/real-time).
  • Familiarity with data governance and lineage tools (e.g., Collibra, BigID, Manta).
  • Knowledge of CI/CD, Git, Jenkins, and code promotion strategies.
  • Strong verbal and written communication skills; experience working in agile/scrum delivery model.

Preferred Qualifications

  • IBM DataStage CP4D certification
  • AWS Certified Data Analytics or Solutions Architect - Associate
  • Snowflake SnowPro certification
  • Experience working in regulated industries (e.g., financial services, utilities, healthcare)

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on jobs.intraedge.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

1:02 min

Applying an ETL methodology to infrastructure configuration management

Axel Barbier · WWC 2023

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

57 sec

Extracting API schemas automatically during continuous integration builds

Axel Barbier · WWC 2023

Videos

See all

Related articles

See all