AWS Data Engineer

Ztek Consulting
United States
11 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours
Job source

Tech stack

Microsoft Access Application Programming Interfaces (APIs) Airflow Amazon Web Services Amazon Elastic Compute Cloud Amazon S3 Microsoft Azure Cloud Computing Computer Programming Databases Continuous Integration Information Engineering
+54 more
Data Governance Data Infrastructure Extract Transform Load (ETL) Data Masking Data Transformation Data Warehousing Software Debugging Software Design Patterns Dimensional Modeling Distributed Computing Environment Github Identity and Access Management JSON Job Scheduling Python (Programming Language) Operational Databases Performance Tuning Query Optimization Power BI Cloud Services SQL Databases Data Streaming Tableau (Software) Unstructured Data Workflow Management Systems Parquet Azure Data Factory Sql Optimization System Availability Git Cloudformation Containerization Data Lakes Pyspark Kubernetes Infrastructure Automation Frameworks Information Technology Avro Data Analytics Enterprise Integration Apache Kafka Spark Streaming Data Management Machine Learning Operations Tools for Reporting Video Streaming Software Coding Terraform Looker Analytics Software Version Control Data Pipelines Docker Jenkins Databricks

Job description

Design, develop, and maintain robust, scalable, and efficient ETL/ELT data pipelines using Databricks, PySpark, and SQL. Build and orchestrate complex data workflows using Apache Airflow, including DAG design, scheduling, monitoring, and error handling. Work extensively within the Databricks on AWS ecosystem (Delta Lake, Unity Catalog, Databricks Workflows, cluster/job optimization); Azure Databricks experience is a plus. Write clean, efficient, and reusable Python code for data transformation, automation, and pipeline development. Optimize SQL queries and data models for performance, scalability, and cost efficiency. Design and implement data lakehouse architectures leveraging Delta Lake best practices (schema evolution, partitioning, Z-ordering, vacuuming, etc.). Integrate data from multiple sources (APIs, databases, streaming platforms, third-party systems) into centralized data platforms. Ensure data quality, integrity, and governance through validation frameworks, monitoring, and alerting. Collaborate closely with Data Analysts, Data Scientists, and Business stakeholders to understand data requirements and deliver reliable datasets. Implement and maintain CI/CD pipelines for data engineering workflows (e.g., using Git, Jenkins, GitHub Actions, or similar). Monitor and troubleshoot production data pipelines, ensuring high availability and minimal downtime. Contribute to architectural decisions around cloud infrastructure, cost optimization, and data platform scalability. Document technical designs, data flows, and operational runbooks. Mentor junior data engineers and contribute to best practices, coding standards, and design patterns within the team.

Requirements

We are looking for an experienced Data Engineer with 5 7 years of hands-on experience building and optimizing large-scale data pipelines and platforms. The ideal candidate has deep expertise in Databricks on AWS (Azure Databricks experience also considered), along with strong skills in Apache Airflow, Python, PySpark, and SQL. You will play a key role in designing, building, and maintaining scalable data infrastructure that powers analytics, reporting, and data science initiatives across the organization., 5 7 years of overall experience in Data Engineering roles. Strong hands-on experience with Databricks (AWS preferred; Azure Databricks acceptable) including Delta Lake, cluster management, job scheduling, and notebook-based development. Proficiency in Apache Airflow for workflow orchestration DAG authoring, sensors, operators, and custom plugins. Strong programming skills in Python, with experience writing production-grade, modular, and testable code. Deep expertise in PySpark for distributed data processing, including performance tuning and optimization techniques. Advanced SQL skills complex joins, window functions, query optimization, and data modeling (dimensional modeling, star/snowflake schemas). Solid understanding of AWS cloud services relevant to data engineering (S3, IAM, EC2, Glue, Lambda, EMR, Redshift, etc.); Azure equivalents (ADLS, ADF, Synapse) a plus. Experience with Delta Lake concepts ACID transactions, time travel, schema enforcement/evolution. Familiarity with version control (Git) and CI/CD practices for data pipelines. Understanding of data warehousing concepts, data lake architectures, and modern lakehouse paradigms. Experience working with structured, semi-structured, and unstructured data (JSON, Parquet, Avro, CSV, etc.). Strong debugging, performance tuning, and problem-solving skills in distributed data processing environments. Good understanding of data governance, security, and compliance practices (role-based access, data masking, encryption). Good to Have

Experience with streaming technologies (Kafka, Kinesis, Spark Structured Streaming). Exposure to Unity Catalog for data governance in Databricks. Knowledge of infrastructure-as-code tools like Terraform or CloudFormation. Experience with containerization (Docker) and orchestration (Kubernetes). Familiarity with BI/reporting tools (Power BI, Tableau, Looker) and how they consume engineered datasets. Exposure to MLOps or Data Science pipeline integration. Relevant certifications: Databricks Certified Data Engineer, AWS Certified Data Analytics/Solutions Architect, or Azure Data Engineer Associate. Educational Qualification

Bachelor’s or Master’s degree in Computer Science, Information Technology, Engineering, or a related field (or equivalent practical experience). Soft Skills

Strong analytical and problem-solving mindset. Excellent communication skills to collaborate with cross-functional teams. Ability to work independently and manage multiple priorities in a fast-paced environment. Detail-oriented with a strong sense of ownership over data quality and pipeline reliability

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · WWC Europe 2026

3:02 min

Audience Q&A on data formats and engine tradeoffs

Matthias Niehoff Matthias Niehoff · WWC Europe 2026

3:47 min

Exploring JSON, CBOR, and JOSE for data serialization

Aaron Russell · LIVE

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

1:52 min

Customizing block storage tiers and formats

Ricardo Sueiras Sueiras · LIVE

2:03 min

Distinguishing type definition constructs from data validation routines

Clemens Vasters Clemens Vasters · WWC 2025

Videos

See all

Related articles

See all