IT - Technology Architect | ETL & Data Quality | ETL & Data Quality - ALL

Cloudspace LLC
Houston, TX, United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Compensation
$124,800.0
Working hours
Regular working hours

Tech stack

Unit Testing Microsoft Azure Big Data Computer Programming Continuous Integration Information Engineering Extract Transform Load (ETL) Data Transformation Python (Programming Language) Query Optimization Azure Data Lake SQL Databases
+12 more
Data Streaming Parquet Data Storage Technologies Azure Data Factory Git Modularization Pytest Data Lakes Pyspark Git Flow Software Coding Databricks

Requirements

Experience: 5 8 years in data engineering with at least 3+ years on Azure. Azure Data Factory: Pipelines, data flows (mapping/wrangling), triggers, integration runtimes, parameterization, error handling & retries. Azure Databricks: Databricks Notebooks, Jobs/Workflows, PySpark/Scala, Delta Lake, cluster configuration and optimization. Programming: Advanced Python or Scala (at least one), strong coding standards, modularization, and unit testing (e.g., pytest). SQL: Expert-level SQL (CTEs, window functions, query tuning), working with large datasets. Data Storage & Formats: ADLS Gen2, Delta Lake, Parquet, partitioning, Z-Ordering, OPTIMIZE/VACUUM. Ops & CI/CD: Git, branching strategies, CI/CD for ADF/Databricks; environment promotion; Infrastructure-as-Code exposure is a plus.

Benefits & conditions

Minimum years of experience: >10 years Certifications Needed: No Top 3 responsibilities you would expect the Subcon to shoulder and execute: Coding AND Development Migration activities Unit testing Interview Process (Is face to face required?) No, Project Code :GPTI core India

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:08 min

Applying software engineering environments and testing to data pipelines

Matthias Niehoff Matthias Niehoff · WWC 2024

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

3:05 min

Tagging and organizing execution scenarios with pytest markers

Florian Bruhin · WWC 2021

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · WWC Europe 2026

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

Videos

See all

Related articles

See all