> Markdown version of [/jobs/ext/2735686-python-spark-developer](https://www.wearedevelopers.com/jobs/ext/2735686-python-spark-developer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Python/Spark Developer - **Company:** Northern Base - **Location:** Pittsburgh, PA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Big Data, Extract Transform Load (ETL), Data Transformation, Relational Databases, Linux, Distributed Systems, Graph Database, Apache Hadoop, Hadoop Distributed File System, Apache Hive, Python (Programming Language), Meta-Data Management, Performance Tuning, SQL Databases, Unstructured Data, Data Processing, Apache Spark, Git, Pyspark, Information Technology, Apache Kafka, Software Version Control, Data Pipelines - **Published:** September 5, 2026 - **Apply:** https://www.careerjet.com/jobad/us3da216b84a7138dda2eb8d0e00b9c4b8 ## About the Role Visa Type: USC / GC / GC EAD Only Must-Have Skills: * Python Development * Apache Spark & PySpark * Big Data Ecosystems Hadoop, Hive, HDFS, Kafka * SQL & Relational Databases * Distributed Computing * Spark Performance Tuning & Optimization * ETL/ELT Pipeline Development * Data Transformation, Cleansing & Validation * Git / Version Control * Linux / Unix * Production Support & Troubleshooting * GenAI Concepts & Applications * AI-Driven Solution Development Additional Skills: * Data Products * Metadata Management * Knowledge Graphs * Ontology * Semantic Technologies, Education: Bachelor's Degree in Computer Science ## Description * Develop and maintain data processing applications using Python and PySpark * Design, build and optimize ETL/ELT pipelines for large-scale datasets * Process structured, semi-structured and unstructured data from multiple sources * Implement data transformation, cleansing and validation frameworks * Collaborate with Data Engineers, Data Scientists, Business Analysts and stakeholders * Optimize Spark jobs for performance, scalability and reliability * Develop reusable data processing components and frameworks * Monitor and troubleshoot production Spark application issues * Identify opportunities to leverage AI and implement AI-driven solutions ## Related Videos - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) ## Related Articles - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [The 13 Best Python Libraries for Developers in 2025](https://www.wearedevelopers.com/magazine/371-the-13-best-python-libraries-for-developers-in-2025) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)