AWS Databricks Engineer

Virtualan Software LLC
Chicago, IL, United States
1 day ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Airflow Amazon Web Services Amazon S3 Cloud Database Databases Continuous Integration Data Architecture Information Engineering Data Governance Extract Transform Load (ETL) Data Systems
+25 more
Data Warehousing Identity and Access Management Python (Programming Language) Performance Tuning Cloud Services SQL Databases Management of Software Versions Data Ingestion Apache Spark Software Troubleshooting AWS Lambda Git Data Lakes Pyspark Integration Tests Information Technology Deployment Automation AWS Glue Apache Kafka Video Streaming Terraform Data Pipelines Amazon Elastic Mapreduce (EMR) Amazon Redshift Databricks

Job description

We are seeking an experienced AWS Databricks Engineer to design, develop, and maintain scalable data engineering solutions using Databricks, Apache Spark, and AWS cloud services. The ideal candidate will have strong experience in data pipelines, ETL/ELT, data lake architecture, and cloud-based data processing., * Design, develop, and maintain scalable data pipelines using Databricks and Apache Spark.

  • Build and optimize ETL/ELT workflows for batch and streaming data processing.
  • Develop solutions using Databricks notebooks, Delta Lake, PySpark, and SQL.
  • Implement and manage data solutions on AWS, including S3, Glue, Lambda, EMR, Redshift, and related services.
  • Develop and maintain data lake/lakehouse architectures using Databricks and AWS.
  • Perform data ingestion from databases, APIs, files, and other enterprise data sources.
  • Implement Delta Lake features such as schema evolution, partitioning, optimization, and data versioning.
  • Monitor, troubleshoot, and optimize data pipelines for performance, reliability, and scalability.
  • Implement security, access controls, data governance, and best practices across AWS and Databricks environments.
  • Work with data architects, analysts, developers, and business stakeholders to understand requirements and deliver data solutions.
  • Implement CI/CD and deployment automation for Databricks and data engineering workloads.
  • Develop unit/integration testing and ensure data quality and pipeline reliability.

Requirements

  • 5+ years of experience in Data Engineering.
  • Strong hands-on experience with Databricks.
  • Strong knowledge of Apache Spark and PySpark.
  • Proficiency in Python and SQL.
  • Strong experience with AWS cloud services, particularly:
  • Amazon S3
  • AWS Glue
  • AWS Lambda
  • Amazon Redshift
  • Amazon EMR
  • IAM
  • Experience with Delta Lake and Lakehouse architecture.
  • Strong understanding of ETL/ELT, data warehousing, and data lake concepts.
  • Experience developing production-grade data pipelines.
  • Experience with Git and CI/CD tools.
  • Strong troubleshooting and performance-tuning skills.

Preferred Skills

  • Experience with Databricks Workflows/Jobs and Unity Catalog.
  • Experience with AWS Step Functions or Airflow.
  • Experience with real-time/streaming technologies such as Kafka or Kinesis.
  • Knowledge of Terraform or Infrastructure as Code.
  • Experience with data governance, lineage, and security.
  • Databricks or AWS certifications are a plus., Bachelor s degree in Computer Science, Information Technology, Engineering, or a related field preferred.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

3:05 min

Audience questions on AI agents and pipeline vectorization

Joy Joy · World Congress 2024

Videos

See all

Related articles

See all