> Markdown version of [/jobs/ext/264296-databricks-data-engineer](https://www.wearedevelopers.com/jobs/ext/264296-databricks-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Databricks Data Engineer - **Company:** Alephys LLC - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $75,000.0 - $140,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Amazon S3, Microsoft Azure, Cloud Computing, Cloud Engineering, Continuous Integration, Data Architecture, Information Engineering, Data Governance, Data Infrastructure, Extract Transform Load (ETL), IBM InfoSphere DataStage, DevOps, Distributed Systems, Github, Apache Hive, Python (Programming Language), Performance Tuning, DataOps, Data Processing, Data Ingestion, Apache Spark, Git, Pyspark, Gitlab-ci, Semi-structured Data, Data Lineage, Terraform, Data Pipelines, Jenkins, Databricks - **Published:** May 14, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=6d75dcabdd05e31b ## About the Role Do you have experience in Spark?, * Experience: 5+ years of experience in Data Engineering, with a heavy focus on the Databricks Data Intelligence Platform. * Core Languages: Expert-level proficiency in Python (PySpark) and advanced Spark SQL. * Databricks Ecosystem: Deep hands-on experience with modern Databricks features, specifically Spark Declarative Pipelines and Unity Catalog for centralized governance. * DevOps / DataOps: Proven ability to implement CI/CD pipelines for data assets using Declarative Asset Bundles (DABs), Git, and automation tools (e.g., GitHub Actions, Jenkins, or GitLab CI). * Architecture & Migration: Strong understanding of distributed computing principles and experience migrating legacy on-premise ETL logic (e.g., DataStage, Informatica) to cloud-native Spark environments. * Problem Solving: Strong analytical skills with the ability to troubleshoot complex data processing issues and optimize massive data transformations. Nice-to-Haves * Experience with event-driven data ingestion architectures and storage-layer triggers (e.g., S3). * Familiarity with cloud infrastructure (AWS, Azure, or GCP) and infrastructure-as-code (Terraform). * Experience building and designing internal frameworks to accelerate data pipeline development. ## Description We are seeking a highly skilled Databricks Data Engineer to design, build, and optimize our next-generation data architecture. In this role, you will be instrumental in modernizing our data infrastructure, transitioning legacy ETL workloads to highly scalable, cloud-native solutions, and ensuring our data assets are secure, discoverable, and performant. You will leverage the latest capabilities of the Databricks platform-from declarative pipelines to advanced governance-to deliver high-quality data products to the business., * Pipeline Engineering: Design, build, and maintain robust, scalable ETL/ELT pipelines using PySpark and Spark SQL to process large volumes of structured and semi-structured data. * Declarative Frameworks: Implement and manage Spark Declarative Pipelines (Delta Live Tables) to simplify pipeline development, automate data quality checks, and streamline operations. * Legacy Modernization: Lead the migration of complex workloads from legacy enterprise ETL systems into modernized, scalable PySpark architectures. * Data Governance & Security: Architect and enforce centralized data governance, access controls, auditing, and data lineage tracking across the organization utilizing Unity Catalog. * CI/CD & Automation: Automate the deployment lifecycle of data pipelines, notebooks, and infrastructure using Databricks Asset Bundles (DABs) integrated with enterprise CI/CD workflows. * Performance Optimization: Profile, tune, and optimize complex PySpark jobs and Spark SQL queries for maximum performance and cost-efficiency. ## Related Videos - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Data Analyst Salary in the UK](https://www.wearedevelopers.com/magazine/278-data-analyst-salary-in-the-uk)