Data Engineer, Data Lake
Role details
Job location
Tech stack
Job description
Note: Visa sponsorship is not available for this position. Candidates must have valid work authorization in the hiring country.
- Design, build, and maintain scalable data pipelines for batch and real-time data processing.
- Develop, optimize, and manage enterprise Data Lake solutions on cloud or on-premises platforms.
- Ingest, transform, and integrate structured, semi-structured, and unstructured data from multiple sources.
- Implement ETL/ELT workflows using modern data engineering tools and frameworks.
- Ensure data quality, governance, security, and compliance across the Data Lake environment.
- Design and optimize data models to support analytics, reporting, and machine learning workloads.
- Monitor, troubleshoot, and improve the performance and reliability of data pipelines.
- Collaborate with Data Scientists, BI Developers, Analysts, and Software Engineers to understand data requirements.
- Implement metadata management, data cataloging, and lineage tracking.
- Automate deployment and infrastructure using CI/CD pipelines and Infrastructure as Code (IaC) where applicable.
- Optimize storage formats (e.g., Parquet, Delta Lake, Iceberg, ORC) and partitioning strategies for performance and cost.
- Work with cloud-native data services such as AWS, Azure, or Google Cloud data platforms.
- Maintain documentation for data architecture, pipelines, and operational procedures.
Requirements
-
Strong experience with Data Lake architectures and modern data platforms.
-
Proficiency in SQL and Python (or Scala/Java).
-
Experience with ETL/ELT frameworks such as Apache Spark, Databricks, Azure Data Factory, AWS Glue, Apache Airflow, or Informatica.
-
Experience with cloud platforms (AWS, Azure, or GCP).
-
Knowledge of data warehousing concepts and dimensional modeling.
-
Familiarity with version control (Git) and CI/CD practices.
-
Understanding of data governance, security, and access control.
-
Experience with Delta Lake, Apache Iceberg, or Apache Hudi.
-
Knowledge of streaming technologies such as Kafka, Kinesis, or Event Hubs.
-
Experience with containerization (Docker, Kubernetes).
-
Cloud certifications (AWS, Azure, or GCP) are a plus.