Databricks Data Engineer

Talent To Hire Inc.
Madrid, Spain
11 days ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience required
6 years minimum
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Big Data BigQuery Cloud Database Cloud Storage Databases Continuous Integration Data Governance Extract Transform Load (ETL) Data Transformation Data Systems Relational Databases
+21 more
Distributed Computing Environment High-Level Architecture Python (Programming Language) Operational Databases Performance Tuning SQL Databases Data Ingestion Sql Optimization Snowflake Apache Spark Git Pyspark Kubernetes Apache Kafka Cloud Optimization Cloudwatch Terraform Data Pipelines Docker Jenkins Databricks

Job description

  • Design, develop, and maintain scalable Databricks-based data pipelines.
  • Build end-to-end data engineering solutions from source ingestion through Gold-layer datasets.
  • Develop high-performance data processing workflows using PySpark, Python, Spark and SQL.
  • Implement Medallion Architecture (Bronze Silver Gold) for ingestion, transformation, aggregation and consumption.
  • Integrate data from multiple sources, including APIs, databases, files, cloud storage and external platforms.
  • Design and optimize Databricks architectures for data ingestion, transformation, processing and storage.
  • Work with large-scale datasets and distributed data processing environments.
  • Perform Spark/Databricks performance tuning to improve processing speed, scalability and cost efficiency.
  • Build robust ETL/ELT workflows with appropriate data quality, monitoring and governance controls.
  • Optimize Databricks workflows, jobs and compute resources.
  • Troubleshoot pipeline performance, reliability and data-quality issues.
  • Collaborate with Data Architects, Data Engineers and business stakeholders to translate requirements into production-ready data solutions.
  • Contribute to engineering standards and Databricks best practices., * Databricks
  • Apache Spark / PySpark
  • Python
  • Advanced SQL
  • Medallion Architecture - Bronze, Silver and Gold
  • End-to-end ETL/ELT data pipeline development
  • Data ingestion from APIs, databases, files and multiple source systems
  • Large-scale/distributed data processing
  • Data transformation and aggregation
  • Databricks/Spark performance tuning and optimization
  • Data quality and pipeline monitoring
  • Cloud-based data engineering environments

Highly Desirable

Experience with some of the following would be advantageous:

  • AWS: S3, Glue ETL, Lambda, Step Functions, ECS, CloudWatch
  • Azure Databricks
  • Snowflake
  • DBT
  • Terraform
  • BigQuery
  • Kafka
  • Docker / Kubernetes
  • Git / Jenkins / CI/CD
  • Data governance
  • Cost monitoring and cloud optimization
  • Infrastructure as Code

Databricks Certified Data Engineer Associate or similar Databricks certification is considered an asset., Describe a PySpark pipeline you developed for a large dataset. What transformations did you implement, and how did you optimize its performance?

  1. Performance: A Databricks/Spark job that previously completed in 20 minutes now takes 90 minutes. How would you diagnose and optimize it?

  2. Data Ingestion: How have you ingested data from APIs, relational databases, files or cloud storage into Databricks?

  3. Data Quality: How do you implement data-quality validation, error handling, monitoring and recovery within a production data pipeline?

  4. Optimization: Give an example where you reduced Databricks/cloud processing costs or significantly improved pipeline performance.

Important: We are prioritizing hands-on engineers, not candidates whose recent experience is primarily management, coordination or high-level architecture.

Requirements

We are looking for a highly hands-on Senior Databricks Data Engineer to design, build, and optimize scalable end-to-end data pipelines.

This role is best suited to an engineer who is comfortable writing PySpark/Python code independently, working with large datasets, integrating multiple data sources, and taking data from initial ingestion through Bronze, Silver, and Gold layers using Medallion Architecture.

This is not a coordination-only or architecture-only position. We are specifically seeking someone who remains hands-on with Databricks, Spark, PySpark, Python and SQL., You are a strong fit if you have personally designed and coded production data pipelines rather than primarily managing other engineers.

You should be able to clearly explain a recent project where you:

Source systems ingestion Bronze Silver Gold downstream consumption

and describe the PySpark/Python code, transformations, architecture decisions, performance improvements and data-quality controls you personally implemented.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:05 min

Enhancing Databricks tooling for software engineering workflows

Alan Mazankiewicz · LIVE

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

Videos

See all

Related articles

See all