> Markdown version of [/jobs/ext/3650680-data-engineer-pyspark](https://www.wearedevelopers.com/jobs/ext/3650680-data-engineer-pyspark). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer (PySpark) - **Company:** Synechron Inc - **Location:** Charlotte, NC, United States - **Salary:** $90,000.0 - $95,000.0 - **Contract:** Permanent contract - **Skills:** Apache HTTP Server, Audit Trail, Automation of Tests, Big Data, BigQuery, Code Review, Continuous Integration, Data Control, Data Discovery, Data Governance, Data Transformation, Data Security, Database Queries, Metadata, Meta-Data Management, Operational Databases, Performance Tuning, Cloudera, Simple Data Format, Software Engineering, Data Streaming, Parquet, Data Ingestion, Apache Spark, Pyspark, Data Lineage, Software Version Control, Data Pipelines - **Published:** October 9, 2026 - **Apply:** https://www.disabledperson.com/jobs/75864990-data-engineer-pyspark ## About the Role * Hands-on experience building and supporting data engineering solutions in production environments. * Strong SQL skills and experience developing data transformations at scale. * Practical experience with Apache Spark, preferably using PySpark. * Experience working with big data file formats, particularly Parquet; familiarity with Iceberg concepts is preferred. * Experience delivering reliable, scheduled data pipelines with monitoring, alerting, and incident response. * Understanding of data cataloging, metadata management, and data lineage concepts, including their importance for governance, discoverability, and reuse. * Exposure to GCP data services such as Dataproc/Spark and Big Query * Understanding of production data operations, reliability, and recoverability. * Solid software engineering practices, including Version control, CI/CD, Automated testing, Code reviews, technical documentation * Ability to collaborate effectively with central cloud, platform, security, and governance teams. * Strong written and verbal communication skills. Preferred, but not required: * Experience building reusable data pipeline frameworks, platform components, or shared engineering capabilities adopted by multiple teams. * Familiarity with Starburst or Trino for federated querying across hybrid environments. * Knowledge of data quality frameworks, reconciliation processes, auditability, and data controls in regulated or enterprise environments. ## Description We are seeking a Data Engineer to design, build, and operate reliable, scalable data solutions in a production environment. This role will focus on developing high-quality data pipelines, extending reusable platform capabilities, and enabling federated teams to onboard and manage data consistently., * Design, build, and harden production-grade data pipelines supporting Data ingestion and curation, Reconciliation and auditability, Data quality monitoring, Metadata and lineage, Notifications and operational alerting * Contribute to and extend the organization's generic pipeline framework, enabling federated teams to onboard and manage data consistently and efficiently. * Build scalable data transformations primarily using PySpark on GCP Dataproc/Spark, with BigQuery as the primary query engine. * Implement best practices for Schema management and evolution, Performance optimization, Reliability and recoverability, Resource utilization and cost efficiency * Work with large-scale datasets and modern data formats, including Parquet and Apache Iceberg. * Integrate data pipelines with centralized data cataloging and governance platforms, such as Dataplex, Knowledge Catalog, or equivalent technologies. * Support automated metadata and lineage harvesting to improve data discovery, governance, and reuse. * Collaborate with platform, security, and central cloud teams to deliver secure, maintainable, and scalable data solutions. * Support hybrid data access and federated query patterns using technologies such as Starburst/Trino, where appropriate. * Develop resilient, observable, and recoverable data flows that meet production service-level agreements. * Participate in production support, incident response, root-cause analysis, and continuous improvement activities. * Apply sound software engineering practices, including code reviews, automated testing, CI/CD, documentation, and version control.