> Markdown version of [/jobs/ext/2578420-data-engineer](https://www.wearedevelopers.com/jobs/ext/2578420-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** Quinnox Inc - **Location:** New York, NY, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Query Performance, Application Programming Interfaces (APIs), Airflow, Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, Big Data, Cloud Computing, Continuous Integration, Data Validation, Data Discovery, Information Engineering, Data Governance, Extract Transform Load (ETL), Data Transformation, Software Debugging, DevOps, Distributed Computing Environment, Github, Apache Hive, Identity and Access Management, JSON, Job Scheduling, Python (Programming Language), Meta-Data Management, Operational Databases, Performance Tuning, Standard Sql, Data Streaming, Parquet, Data Processing, Cloud Platform System, Apache Spark, Caching, Change Data Capture, Gitlab, Git, Data Layers, Data Lakes, Pyspark, Information Technology, Data Lineage, Optimization Algorithms, Deployment Automation, Avro, Software Version Control, Data Pipelines, Databricks - **Published:** August 28, 2026 - **Apply:** https://www.dice.com/job-detail/d7a3e76c-8b36-4b92-a091-2b6b27155bd1 ## About the Role • Hands-on development experience with Databricks notebooks and workflows • Proficiency in Python and PySpark for data transformation and processing • Working knowledge of Unity Catalog for data discovery and lineage • Experience with cluster configuration and job scheduling • Delta Lake development and optimization techniques • Databricks SQL for data analysis and reporting Data Engineering & Pipeline Development • Advanced ETL/ELT pipeline design and development • Delta Lake performance tuning (Z-ordering, data skipping, compaction, vacuuming) • Real-time streaming data pipelines using Structured Streaming and Delta Live Tables • Query performance optimization and debugging slow-running jobs • Data quality validation and testing frameworks • Incremental data processing patterns (CDC, SCD Type 2) Data Processing & Optimization • Spark optimization techniques (partitioning, bucketing, caching, broadcast joins) • Working with large-scale datasets (terabytes to petabytes) • Data pipeline orchestration and scheduling • Monitoring and alerting data pipelines • Implementing Bronze/Silver/Gold (Medallion) data layer patterns Required Technical Skills Databricks & Spark Proficiency: 3+ years of hands-on experience building data pipelines in Databricks; deep understanding of Spark fundamentals, transformations, actions, and performance optimization techniques including partitioning, caching, and resource management Advanced PySpark and SQL Skills: Expert-level proficiency writing production-quality PySpark code and complex SQL queries for data transformation, aggregation, and analysis; experience with DataFrame API, Spark SQL, and UDFs; strong understanding of lazy evaluation and execution plans Data Engineering & ETL/ELT: Proven experience building and maintaining production data pipelines; hands-on experience with incremental data loading, change data capture (CDC), and slowly changing dimensions; experience handling data quality issues and implementing data validation frameworks Cloud & Big Data Technologies: Strong proficiency with AWS services (S3, EC2, IAM, Glue, Athena); experience working with large-scale distributed data processing; familiarity with data formats (Parquet, Delta, JSON, Avro) and compression techniques DevOps & CI/CD: Experience with version control (Git) and CI/CD pipelines using GitLab, GitHub Actions, or similar tools; familiarity with testing data pipelines and deployment automation; experience with Databricks Repos and workspace-level integrations Data Governance: Understanding of data lineage, cataloging, and metadata management; experience implementing data quality checks and monitoring; knowledge of data privacy and security best practices in cloud environments (nice to have) Preferred Qualifications • Bachelor's degree in Computer Science, Engineering, or related field • Databricks Certified Data Engineer Associate or Professional certification • Experience with data orchestration tools (Apache Airflow, Databricks Workflows) • Strong debugging and problem-solving skills • Excellent communication skills for technical collaboration ## Description We are seeking an experienced Data Engineer to join our team and build robust, scalable data pipelines. In this role, you will: • Design and implement scalable PySpark data pipelines for batch and streaming workloads • Optimize Spark jobs and queries for performance and cost efficiency • Build and maintain ETL/ELT processes following data engineering best practices • Troubleshoot and resolve complex data pipeline and processing issues • Collaborate with data teams to ensure data quality and reliability ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Tips and Tricks for Working with JSON](https://www.wearedevelopers.com/videos/1229-tips-and-tricks-for-working-with-json) - [From event streaming to event sourcing 101](https://www.wearedevelopers.com/videos/91-from-event-streaming-to-event-sourcing-101) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Introducing JSON Structure](https://www.wearedevelopers.com/videos/100219-introducing-json-structure) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) ## Related Articles - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again)