Data Engineer GCP AWS & Databricks

iShare Inc
Ontario, CA, United States
12 days ago
Apply on www.wayup.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Agile Methodology Airflow Amazon Web Services Amazon S3 Big Data BigQuery Cloud Database Cloud Storage Computer Programming Continuous Integration Data Architecture Information Engineering
+24 more
Data Integration Extract Transform Load (ETL) Data Security Data Warehousing Data Flow Control Python (Programming Language) Cloud Services Cloudera SQL Databases Unstructured Data Workflow Management Systems Google Cloud Apache Spark Git Data Lakes Pyspark Infrastructure Automation Frameworks Real Time Data Apache Kafka Data Management Video Streaming Terraform Data Pipelines Databricks

Job description

We are seeking an experienced Data Engineer to design, develop, and maintain scalable cloud-based data pipelines and data platforms. The ideal candidate will have strong hands-on experience with GCP, AWS, Databricks, Apache Spark, PySpark, Python, and SQL. This role requires expertise in building reliable ETL/ELT pipelines, integrating data from multiple sources, and optimizing cloud data solutions for performance, security, scalability, and cost., Design, develop, and maintain scalable ETL/ELT data pipelines. Build cloud-based data solutions using GCP and AWS services. Use Databricks, Apache Spark, and PySpark for large-scale data processing. Integrate structured, semi-structured, and unstructured data from multiple sources. Develop and optimize batch and real-time data-processing workflows. Improve pipeline performance, reliability, scalability, and cost efficiency. Implement data-quality checks, monitoring, security, and governance standards. Design and support cloud data warehouses and data lakes. Troubleshoot production issues and perform root-cause analysis. Collaborate with data architects, analysts, application teams, and business stakeholders. Create and maintain technical documentation for pipelines, data models, and workflows.

Requirements

5+ years of professional data engineering experience. Strong hands-on experience with both GCP and AWS. Expertise in Databricks, Apache Spark, and PySpark. Strong programming skills in Python and SQL. Proven experience developing ETL/ELT pipelines and cloud data platforms. Experience with data warehouses, data lakes, and dimensional data modeling. Experience with orchestration tools such as Apache Airflow or Google Cloud Composer. Understanding of data security, governance, monitoring, and quality frameworks. Strong analytical, troubleshooting, and problem-solving skills. Excellent communication and cross-functional collaboration skills. Preferred Qualifications Experience with GCP services such as BigQuery, Cloud Storage, Dataflow, Dataproc, Pub/Sub, and Cloud Composer. Experience with AWS services such as S3, Glue, EMR, Redshift, Lambda, and Kinesis. Familiarity with Delta Lake and Databricks Lakehouse architecture. Experience with streaming technologies such as Apache Kafka. Familiarity with CI/CD, Git, Terraform, and cloud infrastructure automation. Experience working in Agile development environments. Must-Have Skills GCP | AWS | Databricks | Apache Spark | PySpark | Python | SQL | ETL/ELT | Airflow/Cloud Composer | Data Warehousing | Data Lakes

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.wayup.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

3:05 min

Audience questions on AI agents and pipeline vectorization

Joy Joy · World Congress 2024

Videos

See all

Related articles

See all