> Markdown version of [/jobs/ext/2217568-pyspark-developer](https://www.wearedevelopers.com/jobs/ext/2217568-pyspark-developer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # PySpark Developer - **Company:** RAVIN IT SOLUTIONS, Inc - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Agile Methodology, Airflow, Amazon Web Services, Amazon S3, Apache HTTP Server, Application Frameworks, Big Data, Information Systems, Continuous Integration, Information Engineering, Data Governance, Data Infrastructure, Extract Transform Load (ETL), Data Transformation, Data Systems, DevOps, Distributed Computing Environment, Identity and Access Management, Python (Programming Language), Metadata, Performance Tuning, Query Optimization, Cloud Services, Cloudera, Software Engineering, SQL Databases, Unstructured Data, Data Processing, Cloud Platform System, Apache Spark, Git, Amazon Relational Database Service, Data Lakes, Pyspark, Information Technology, Integration Frameworks, Apache Kafka, Cloudwatch, Data Pipelines - **Published:** August 25, 2026 - **Apply:** https://www.dice.com/job-detail/801a2add-9b80-4a88-b697-8d628623b308 ## About the Role * Bachelor''s degree in Computer Science, Information Systems, Engineering, or related field. * 10+ years of IT experience with 6+ years in Data Engineering and Big Data development. * Strong hands-on experience with PySpark and Spark-based data processing. * Experience developing data pipelines on Cloudera Data Platform (CDP). * Strong knowledge of Apache Iceberg table architecture and data lake concepts. * Experience designing and managing workflows using Apache Airflow. * Experience with SQL and data modeling techniques. * Strong understanding of ETL/ELT development and data transformation frameworks. * Experience working with large-scale structured and unstructured datasets. * Knowledge of Git, CI/CD pipelines, and DevOps practices. * Strong analytical, troubleshooting, and problem-solving skills. * Excellent communication and collaboration skills. Preferred Qualifications * Experience with AWS services including S3, Glue, Lambda, ECS/EKS, EMR, RDS, IAM, and CloudWatch. * Experience building cloud-native data lake and Lakehouse solutions. * Strong knowledge of data partitioning, performance tuning, and query optimization. * Experience with Kafka or event-driven data processing frameworks. * Experience with Java or Python application development. * Experience implementing data quality, metadata, and data governance solutions. * Healthcare, Medicaid, Claims Processing, or Insurance industry experience. * Experience working in Agile development environments. * AWS Certification and/or Cloudera Certification preferred. * Experience supporting large-scale cloud migration and modernization initiatives. ## Description We are seeking a highly skilled Senior PySpark Developer to support a large-scale cloud transformation initiative focused on building modern data platforms on Cloudera, Apache Iceberg, Airflow, and AWS. The role will be responsible for designing, developing, and optimizing scalable data pipelines that ingest, transform, and process large volumes of enterprise data. The successful candidate will work closely with data architects, business analysts, cloud engineers, and application teams to build reliable, high-performance data solutions supporting analytics, reporting, and operational workloads. Responsibilities include developing PySpark applications, implementing ETL/ELT processes, designing Iceberg-based data models, orchestrating workflows using Airflow, and leveraging AWS services to support cloud-native data processing. This requires strong expertise in distributed data processing, performance tuning, data quality, and cloud-based architectures. The candidate will participate in end-to-end solution delivery, from requirements analysis and solution design through development, testing, deployment, and production support. They will also contribute to establishing best practices, reusable frameworks, CI/CD processes, and data governance standards across the platform. The ideal candidate is a hands-on developer with deep technical expertise in modern data engineering technologies and experience working in large-scale enterprise cloud migration and modernization programs. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) ## Related Articles - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again)