> Markdown version of [/jobs/ext/1091948-data-engineer-python-pyspark-puerto-rico](https://www.wearedevelopers.com/jobs/ext/1091948-data-engineer-python-pyspark-puerto-rico). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer (Python/PySpark) (Puerto Rico) - **Company:** RTX - **Location:** ISABEL, United States - **Contract:** Permanent contract - **Skills:** Airflow, Amazon Web Services, Microsoft Azure, Big Data, BigQuery, Code Review, Directed Acyclic Graph (Directed Graphs), Extract Transform Load (ETL), Data Structures, Distributed Systems, Apache Hive, Python (Programming Language), Machine Learning, Object-Oriented Software Development, Raw Data, Cloud Services, Flask (Web Framework), Snowflake, Concurrency, Apache Spark, Git, Fastapi, Data Lakes, Pyspark, Deployment Automation, Apache Kafka, Spark Streaming, Video Streaming, Restful APIs, Data Pipelines, Docker, Amazon Redshift, Databricks - **Published:** June 30, 2026 - **Apply:** https://globalhr.wd5.myworkdayjobs.com/REC_RTX_Ext_Gateway/job/US-PR-SANTA-ISABEL-B1--Felicia-Industrial-Park---St-B1--BLDG-1/Data-Engineer--Python-PySpark---Puerto-Rico-_01855754 ## About the Role Collins Aerospace is seeking an experienced Python and PySpark Developer to design, build, and optimize our next-generation big data pipelines. In this role, you will handle large-scale datasets, optimize distributed computing clusters, and bridge the gap between raw data ingestion and production-ready analytics. The ideal candidate thrives on optimizing cluster performance, resolving data skewness, and writing clean, maintainable Python code., * Must be a U.S. Citizen. * Strong proficiency in Python (OOP, concurrency, data structures) and advanced SQL. * Deep production experience with Apache Spark / PySpark (Data Frames, Spark SQL, RDDs). * Hands-on experience with cloud data platforms like AWS (EMR, Glue), Azure (Databricks), or GCP. * Experience working with Snowflake, Big Query, Redshift, or Synapse. * Proficient with Git, Docker, and automated deployment pipelines. Qualifications We Prefer: * PySpark MLlib or deploying Machine Learning models to production. * Familiarity with streaming technologies like Apache Kafka or Spark Structured Streaming. Databricks Certified Data Engineer or Apache Spark Developer certifications. ## Description * Design and deploy robust batch and streaming ETL/ELT pipelines using PySpark and Python. * Optimize Spark jobs by tuning configurations, managing partitioning, and resolving data skew or OOM (Out of Memory) errors. * Implement modern data lakehouse architectures using Delta Lake, Iceberg, or Hudi. * Build and maintain complex workflow DAGs using orchestration tools like Apache Airflow. * Develop backend Python services or REST APIs (e.g., Fast API, Flask) to expose processed data to downstream applications. * Write clean, modular, and unit-tested code while participating in rigorous code reviews ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [The 13 Best Python Libraries for Developers in 2025](https://www.wearedevelopers.com/magazine/371-the-13-best-python-libraries-for-developers-in-2025) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)