> Markdown version of [/jobs/ext/1900485-data-software-engineer-azure-databricks](https://www.wearedevelopers.com/jobs/ext/1900485-data-software-engineer-azure-databricks). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Software Engineer, Azure Databricks - **Company:** EPAM Systems, Inc. - **Location:** Newtown, PA, United States (Remote available) - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Agile Methodology, Big Data, Computer Programming, Continuous Integration, Data Architecture, Data Governance, Programming Tools, Python (Programming Language), Performance Tuning, Scrum Methodology, Azure Data Lake, SAP HANA, Software Construction, Workflow Management Systems, Build Management, Data Lakes, Data Programming, Data Management, Data Pipelines, Databricks - **Published:** August 1, 2026 - **Apply:** https://arc.dev/remote-jobs/j/redirect/p8u238xyjs ## About the Role * Strong experience as a Data Engineer with proficiency in Databricks * 3+ years of experience building reusable libraries, SDKs or internal developer tooling * Knowledge of Data Mesh/Data Product concepts including data ownership, domain-oriented design and self-serve data platforms * Deep expertise in Delta Lake, Delta tables and compaction with optimization for high-performance workloads * Proven experience designing and maintaining complex data pipelines on cloud object stores (ADLS, etc.) * Strong programming skills in Python for data engineering workloads * Solid understanding of Lakehouse architecture and best practices for large-scale data platforms * Hands-on experience with data pipeline monitoring, troubleshooting and performance tuning * Familiarity with CI/CD and workflow orchestration (Databricks Jobs) * Experience working in agile teams with a focus on ownership, autonomy and best practices * Excellent problem-solving skills and capability to handle high-scale, complex data challenges * English proficiency at B2 level or higher Nice to have * Experience with agile methodologies (Scrum, SAFe) * Experience developing enablement materials and technical documentation ## Description We are looking for a Data Software Engineer to help build a scalable solution used by data engineers to build, test, deploy and monitor data pipelines. You will develop reusable libraries and SDK modules that enable teams to ship data products autonomously and efficiently, working with Databricks, Delta Lake, Azure Data Lake and SAP HANA Data Lake., * Develop and maintain reusable SDK modules (Python) for data pipeline operations such as compaction, quality checking, schema evolution and table lifecycle management * Design and build framework libraries following software engineering best practices, ensuring scalability and reliability across 100+ pipelines and diverse data domains * Collaborate with Data Engineers who consume the framework to gather feedback, understand pain points and iterate on the SDK * Contribute to platform architecture discussions and performance tuning strategies * Write high-quality documentation and contribute to enablement materials for framework consumers * Help define and infuse data engineering best practices through enablement, SDKs and templates * Ensure data quality, consistency and governance across Lakehouse environments ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Cutting LLM Costs Without Cutting Quality: How to Beat Proprietary LLMs with Fine-Tuned Open Source](https://www.wearedevelopers.com/videos/100151-cutting-llm-costs-without-cutting-quality-how-to-beat-proprietary-llms-with-fine-tuned-open-source) - [Evolving the developer experience in the age of AI](https://www.wearedevelopers.com/videos/100201-evolving-the-developer-experience-in-the-age-of-ai) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Parquet, Delta, Iceberg & Ducklake - An introduction for developers](https://www.wearedevelopers.com/videos/100075-parquet-delta-iceberg-ducklake-an-introduction-for-developers) ## Related Articles - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 132 - Binging WADFlix?](https://www.wearedevelopers.com/magazine/473-dev-digest-132-binging-wadflix)