Data Engineer

VIIS GLOBAL LLC
Pasadena, CA, United States
3 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Working hours
Regular working hours
Job source

Tech stack

Adobe InDesign Application Programming Interfaces (APIs) Airflow Amazon Web Services Amazon S3 Big Data Databases Continuous Integration Data Architecture Data Validation Information Engineering Data Governance
+23 more
Extract Transform Load (ETL) Data Transformation Database Queries Apache Hive Identity and Access Management Python (Programming Language) Web Application Frameworks Data Processing Data Storage Technologies Apache Spark AWS Lambda Git Data Lakes Pyspark AWS Glue AWS Data Analytics Apache Kafka Cloudwatch Terraform Data Pipelines Amazon Elastic Mapreduce (EMR) Amazon Redshift Databricks

Job description

We are seeking an experienced Data Engineer with strong hands-on expertise in AWS, Databricks, PySpark, and Python. The ideal candidate will have experience designing and developing scalable data pipelines and data processing solutions using AWS cloud services and Databricks., * Design, develop, and maintain scalable data pipelines using Databricks, PySpark, Python, and AWS.

  • Develop robust ETL/ELT pipelines for ingesting and transforming large volumes of data.
  • Build and optimize PySpark/Spark SQL jobs within Databricks.
  • Work with AWS S3 for data storage and data lake solutions.
  • Develop data processing workflows using Databricks Workflows and related AWS services.
  • Work with Delta Lake for reliable data storage, incremental processing, and data transformation.
  • Develop reusable Python frameworks and utilities for data engineering processes.
  • Perform data validation, quality checks, error handling, and reconciliation.
  • Troubleshoot production pipeline issues and perform Root Cause Analysis (RCA).
  • Optimize Spark jobs for performance, scalability, and cost efficiency.
  • Integrate data from databases, APIs, files, and other enterprise data sources.
  • Implement CI/CD and source-control practices for data engineering applications.
  • Collaborate with Data Architects, Data Scientists, Analysts, and business stakeholders.
  • Participate in design, development, testing, deployment, and production support.

Requirements

  • 10+ years of overall Data Engineering experience preferred.
  • Strong hands-on experience with AWS.
  • Strong experience with Databricks.
  • Strong hands-on experience with PySpark / Apache Spark.
  • Strong Python programming experience.
  • Strong SQL skills.
  • Experience developing enterprise-scale ETL/ELT pipelines.
  • Experience with AWS S3 and AWS data services.
  • Experience with Delta Lake.
  • Experience working with large-scale datasets.
  • Strong understanding of data lake/lakehouse architecture.
  • Experience with Spark performance tuning and optimization.
  • Experience with Git and CI/CD.

AWS Skills

Candidates should have hands-on experience with AWS services such as:

  • Amazon S3
  • AWS Glue
  • AWS Lambda
  • Amazon Redshift
  • Amazon EMR
  • AWS CloudWatch
  • AWS IAM

Preferred Skills

  • Databricks certification
  • Experience with Unity Catalog
  • Experience with Airflow
  • Experience with Kafka
  • Experience with AWS Glue/Airflow orchestration
  • Experience with Terraform
  • Experience with data governance and data quality frameworks

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

3:05 min

Audience questions on AI agents and pipeline vectorization

Joy Joy · World Congress 2024

Videos

See all

Related articles

See all