Databricks Data Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+21 more
Job description
- Design, develop, and maintain scalable Databricks-based data pipelines.
- Build end-to-end data engineering solutions from source ingestion through Gold-layer datasets.
- Develop high-performance data processing workflows using PySpark, Python, Spark and SQL.
- Implement Medallion Architecture (Bronze Silver Gold) for ingestion, transformation, aggregation and consumption.
- Integrate data from multiple sources, including APIs, databases, files, cloud storage and external platforms.
- Design and optimize Databricks architectures for data ingestion, transformation, processing and storage.
- Work with large-scale datasets and distributed data processing environments.
- Perform Spark/Databricks performance tuning to improve processing speed, scalability and cost efficiency.
- Build robust ETL/ELT workflows with appropriate data quality, monitoring and governance controls.
- Optimize Databricks workflows, jobs and compute resources.
- Troubleshoot pipeline performance, reliability and data-quality issues.
- Collaborate with Data Architects, Data Engineers and business stakeholders to translate requirements into production-ready data solutions.
- Contribute to engineering standards and Databricks best practices., * Databricks
- Apache Spark / PySpark
- Python
- Advanced SQL
- Medallion Architecture - Bronze, Silver and Gold
- End-to-end ETL/ELT data pipeline development
- Data ingestion from APIs, databases, files and multiple source systems
- Large-scale/distributed data processing
- Data transformation and aggregation
- Databricks/Spark performance tuning and optimization
- Data quality and pipeline monitoring
- Cloud-based data engineering environments
Highly Desirable
Experience with some of the following would be advantageous:
- AWS: S3, Glue ETL, Lambda, Step Functions, ECS, CloudWatch
- Azure Databricks
- Snowflake
- DBT
- Terraform
- BigQuery
- Kafka
- Docker / Kubernetes
- Git / Jenkins / CI/CD
- Data governance
- Cost monitoring and cloud optimization
- Infrastructure as Code
Databricks Certified Data Engineer Associate or similar Databricks certification is considered an asset., Describe a PySpark pipeline you developed for a large dataset. What transformations did you implement, and how did you optimize its performance?
-
Performance: A Databricks/Spark job that previously completed in 20 minutes now takes 90 minutes. How would you diagnose and optimize it?
-
Data Ingestion: How have you ingested data from APIs, relational databases, files or cloud storage into Databricks?
-
Data Quality: How do you implement data-quality validation, error handling, monitoring and recovery within a production data pipeline?
-
Optimization: Give an example where you reduced Databricks/cloud processing costs or significantly improved pipeline performance.
Important: We are prioritizing hands-on engineers, not candidates whose recent experience is primarily management, coordination or high-level architecture.
Requirements
We are looking for a highly hands-on Senior Databricks Data Engineer to design, build, and optimize scalable end-to-end data pipelines.
This role is best suited to an engineer who is comfortable writing PySpark/Python code independently, working with large datasets, integrating multiple data sources, and taking data from initial ingestion through Bronze, Silver, and Gold layers using Medallion Architecture.
This is not a coordination-only or architecture-only position. We are specifically seeking someone who remains hands-on with Databricks, Spark, PySpark, Python and SQL., You are a strong fit if you have personally designed and coded production data pipelines rather than primarily managing other engineers.
You should be able to clearly explain a recent project where you:
Source systems ingestion Bronze Silver Gold downstream consumption
and describe the PySpark/Python code, transformations, architecture decisions, performance improvements and data-quality controls you personally implemented.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Making Data Warehouses Fast: A Developer’s Story
Top Big Data Technologies That You Need to Know
Where To Find Software Engineering Jobs
Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production