> Markdown version of [/jobs/ext/1930686-pyspark-data-engineer](https://www.wearedevelopers.com/jobs/ext/1930686-pyspark-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # PySpark Data Engineer - **Company:** DISCLAIMERS LLC - **Location:** New York, NY, United States - **Experience:** Expert - **Salary:** $90,000.0 - $110,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, Cloud Computing, Cloud Database, Databases, Continuous Delivery, Continuous Integration, Information Engineering, Extract Transform Load (ETL), Data Warehousing, Distributed Systems, Python (Programming Language), Machine Learning, Performance Tuning, Systems Development Life Cycle, Systems Architecture, Data Logging, Scripting, Data Ingestion, Snowflake, Apache Spark, Software Troubleshooting, Pyspark, Star Schema, Data Management, Machine Learning Operations, Data Pipelines, Databricks - **Published:** August 5, 2026 - **Apply:** https://www.careerbuilder.com/job-details/pyspark-data-engineer-with-databricks-new-york-ny--e102f177-7173-40b6-8343-b0ac98392112 ## About the Role * 8+ years of experience in data engineering with strong hands-on work in PySpark and Python. * Deep experience with Databricks, Spark optimization, cluster tuning, and performance troubleshooting. * Strong experience working with Snowflake or similar cloud data warehouses. * Practical knowledge of workflow orchestration tools and dependency management. * Solid understanding of data modeling, ingestion frameworks, and distributed systems architecture. * Hands-on experience implementing CI/CD for data and ML pipelines. * Strong experience with MLflow for managing the ML lifecycle. * Excellent communication skills with the ability to work across engineering and business teams., Apache Spark, Artificial Intelligence (AI), Business Transformation, Cloud Computing, Communication Skills, Compensation and Benefits, Continuous Deployment/Delivery, Continuous Integration, Data Management, Data Modeling, Data Quality, Data Science, Data Warehousing, Database Extract Transform and Load (ETL), Distributed Computing, Ecosystems, Employee Assistance Plan, Equal Employment Opportunity (EEO), Healthcare, Identify Issues, International Business, Legal, Machine Learning, Operations Planning, Performance Tuning/Optimization, Production Systems, Python Programming/Scripting Language, Reconciliation, Scalable System Development, Snowflake Schema, System Architecture, Team Player, Technical/Engineering Design, Use Cases ## Description * Design, develop, and maintain end-to-end ETL/ELT pipelines using Python and PySpark on Databricks. * Optimize Spark jobs for performance, scalability, and cost-efficiency in production environments. * Implement data quality frameworks including validation, reconciliation, and anomaly detection. * Build and manage orchestration workflows (Airflow / Databricks Workflows / equivalent). * Implement pipeline monitoring, logging, alerting, and observability for reliable operations. * Develop and operationalize ML workflows using MLflow (experiment tracking, model registry, packaging, deployment). * Build scalable data ingestion and data modeling solutions for analytics and ML use cases. * Collaborate with data scientists, platform teams, engineering stakeholders, and business partners. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [How Cisco embraced a DevOps culture within its network engineering team](https://www.wearedevelopers.com/videos/99-how-cisco-embraced-a-devops-culture-within-its-network-engineering-team) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) ## Related Articles - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)