Big Data Engineer

Wise Skulls llc
Dallas, TX, United States
6 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Big Data Cloudera Impala Code Review Information Engineering Extract Transform Load (ETL) Data Structures Data Systems Distributed Computing Environment Distributed Systems Apache Hadoop Hadoop Distributed File System
+11 more
Apache Hive Job Scheduling Python (Programming Language) Machine Learning SQL Databases Apache Spark Data Management Software Coding Software Version Control Data Pipelines Control M

Job description

  • Lead the design, development, and maintenance of robust, scalable data pipelines for ingestion, transformation, and processing of large datasets in an on-premises environment.
  • Own architectural and design decisions for data solutions, evaluating trade-offs and defining technical standards for the team.
  • Mentor, guide, and support other data engineers through code reviews, design reviews, technical coaching, and hands-on problem-solving.
  • Build and optimize data workflows using Python, Spark, and the Hadoop ecosystem.
  • Work extensively with Hadoop ecosystem components (Hive, HDFS, Impala) to manage and query large-scale data.
  • Manage and optimize batch scheduling and job orchestration using enterprise schedulers such as CA7 or Control-M.
  • Ensure data quality, integrity, and performance across data platforms.
  • Collaborate with data analysts, data scientists, and business stakeholders to translate data requirements into sound technical designs.
  • Troubleshoot and resolve complex issues in data pipelines and production environments, acting as an escalation point for the team.
  • Champion best practices for coding standards, version control, testing, and documentation.
  • Stay current with emerging technologies, particularly AI/ML capabilities, and identify opportunities to apply them to data engineering workflows.

Requirements

We are looking for an experienced Big Data Engineer with 7-10 years of hands-on experience to design, build, and maintain scalable data pipelines and processing systems in an on-premises Big Data environment. Beyond strong individual contribution, the ideal candidate will own architecture and design decisions, set technical direction, and mentor and support other developers on the team. The role works closely with cross-functional teams to deliver reliable, high-quality data solutions that support business and analytics needs., * 7-10 years of overall experience in data engineering, with a proven track record in technical leadership (design ownership, mentoring, guiding development teams).

  • Python - strong hands-on development experience building production-grade data solutions.
  • Big Data / Hadoop ecosystem (Hadoop, Hive, Impala, HDFS) - deep, hands-on experience in on-premises environments.
  • Apache Spark - solid experience developing and tuning large-scale distributed data processing jobs.
  • Job scheduling / orchestration - hands-on experience with CA7 or Control-M (or comparable enterprise schedulers).
  • Strong understanding of data structures, ETL processes, and SQL.
  • Extensive experience with large-scale data processing and distributed systems.
  • Demonstrated ability to make sound architecture/design decisions and to mentor and support other developers.
  • Exposure to AI/ML concepts or tools, with a strong willingness to learn and grow in this space.
  • Banking/Financial domain highly preferred

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:39 min

Addressing code review surrender and process exploitation

Laura Tacho Laura Tacho · WWC Europe 2026

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

56 sec

The hidden costs of delayed peer code reviews

Tim Gilboy Tim Gilboy

2:10 min

Why organizations combine big data and machine learning

Ayon Roy · LIVE

Videos

See all

Related articles

See all