Big Data Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+7 more
Job description
Experteer Overview As a Big Data Engineer at EXL, you will own the design, development, and maintenance of scalable on-prem data pipelines that feed analytics and business insights. You’ll set technical direction, guide a team of engineers, and collaborate with data analysts, scientists, and stakeholders to translate requirements into robust data solutions. You’ll leverage Python, Spark, and the Hadoop ecosystem to build reliable workflows and ensure data quality at scale. This role combines hands-on engineering with architectural ownership to drive high-impact data platforms and stay ahead with emerging AI/ML opportunities. Compensation / Benefits * Lead design, development, and maintenance of scalable data pipelines for ingestion, transformation, and processing of large datasets on-premises * Own architectural decisions and define technical standards for the team * Mentor and support data engineers through reviews, coaching, and hands-on problem-solving * Build and optimize data workflows using Python, Spark, and the Hadoop ecosystem * Work with Hive, HDFS, Impala to manage and query large-scale data * Manage batch scheduling and orchestration with CA7 or Control-M * Ensure data quality, integrity, and performance across platforms * Collaborate with analysts, scientists, and stakeholders to translate requirements into designs * Troubleshoot complex issues and act as escalation point * Champion coding standards, testing, and documentation * Stay current with emerging technologies, especially AI/ML capabilities, and apply where appropriate Tasks * 7-10 years of data engineering experience with proven technical leadership * Strong Python development for production-grade data solutions * Deep on-premises Big Data / Hadoop ecosystem experience (Hadoop, Hive, Impala, HDFS) * Apache Spark experience for large-scale distributed processing * Hands-on experience with CA7 or Control-M (or comparable schedulers) * Solid understanding of data structures, ETL, and SQL * Extensive experience with large-scale data processing and distributed systems * Ability to make sound architecture/design decisions and mentor others * Exposure to AI/ML concepts or tools with willingness to learn Key requirements *
Requirements
proven using Python, Spark, and the Hadoop ecosystem * Work with Hive, HDFS, Impala to manage and query large-scale data * Manage batch scheduling and orchestration with CA7 or Control-M * Ensure data quality, integrity, and performance across platforms * Collaborate with analysts, scientists, and stakeholders to translate requirements into designs * Troubleshoot complex issues and act as escalation point * Champion coding standards, testing, and documentation * Stay current with emerging technologies, especially AI/ML capabilities, and apply where appropriate Tasks * 7-10 years of data engineering experience with proven technical leadership * Strong Python development for production-grade data solutions * Deep on-premises Big Data / Hadoop ecosystem experience (Hadoop, Hive, Impala, HDFS) * Apache Spark experience for large-scale distributed processing * Hands-on experience with CA7 or Control-M (or comparable schedulers) * Solid understanding of data structures, ETL, and SQL * aaaaa data experience with large-scale data processing and distributed systems * Ability to make sound architecture/design decisions and mentor others * Exposure to AI/ML concepts or tools with willingness to learn Key requirements *
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on us.experteer.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Data Engineer Salary UK
How to Become an AI Engineer
Highest Paying Tech Companies for Developers
Making Data Warehouses Fast: A Developer’s Story