Data Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+34 more
Job description
We are looking for an experienced Data Engineer to design, develop, and maintain scalable data pipelines and data processing solutions supporting large-scale enterprise platforms. The ideal candidate will have strong hands-on experience with Python, PySpark, SQL, cloud data platforms, distributed data processing, and data pipeline development.
The Data Engineer will work closely with data scientists, software engineers, product teams, architects, and other stakeholders to build reliable, scalable, and high-performance data solutions., Design, develop, and maintain scalable batch and real-time data pipelines Build ETL/ELT workflows using Python, PySpark, SQL, and cloud technologies Develop distributed data processing solutions using Apache Spark/Databricks Build data pipelines integrating data from multiple sources and systems Develop streaming and event-driven data pipelines using Kafka Implement workflow orchestration using Airflow or similar technologies Design and optimize data models, queries, and data processing jobs Work with cloud data platforms such as Azure and Google Cloud Platform Implement data quality checks, validation, monitoring, and error handling Troubleshoot pipeline failures and performance issues Optimize Spark jobs, SQL queries, and data processing workflows for performance and scalability Collaborate with data scientists, software engineers, product managers, architects, and DevOps teams Participate in technical design, code reviews, testing, deployment, and production support Follow engineering best practices for security, scalability, reliability, and maintainability Preferred Skills
Requirements
5+ years of experience in Data Engineering or a closely related field Strong programming experience with Python Strong hands-on experience with Apache Spark / PySpark Strong expertise in SQL, including complex queries, optimization, and data transformations Experience designing and developing ETL/ELT data pipelines Experience working with Databricks or similar distributed data processing platforms Strong experience with cloud data platforms, preferably Azure or Google Cloud Platform Experience with cloud storage and data services such as Azure Data Lake, ADLS, GCS, BigQuery, or equivalent Experience with Apache Kafka or other event-streaming technologies Experience with Airflow / Cloud Composer or similar workflow orchestration tools Strong understanding of data modeling, data warehousing, and data architecture Experience working with both relational and NoSQL databases Experience with Git and CI/CD Strong understanding of data quality, monitoring, testing, and pipeline reliability, Experience with Databricks Strong experience with PySpark and Python Experience with Scala is a plus Experience with Google Cloud Platform, BigQuery, GCS, Dataproc, and Cloud Composer Experience with Azure, ADLS, Azure Data Factory, and Azure Databricks Experience with Kafka Experience with Airflow Experience with Snowflake Experience with PostgreSQL or other relational databases Experience with NoSQL databases Experience with Terraform, Docker, or Kubernetes Experience developing real-time/streaming data pipelines Experience working with large-scale datasets and high-volume data processing Experience in retail, e-commerce, marketplace, or other large-scale enterprise environments
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Loading talks and stories from around this roleβ¦