Hadoop Hive Python Developer

DTEL Engineering & Consultants Inc
Charlotte, United States
7 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
9 years minimum
Working hours
Regular working hours
Job source

Tech stack

Airflow Data Analysis Big Data Cloud Computing Continuous Integration Data Architecture Data Governance Data Infrastructure Extract Transform Load (ETL) Data Systems Distributed Systems Apache Hadoop
+23 more
Hadoop Distributed File System MapReduce Apache Hive Python (Programming Language) Apache Oozie Scrum Methodology Query Optimization SQL Databases Data Streaming Unstructured Data Workflow Management Systems Cloud Platform System Apache Yarn Delivery Pipeline Git Data Lakes Pyspark Real Time Data Apache Kafka Bitbucket Tez (Software) Jenkins Databricks

Job description

Big Data Platform Engineering

  • Design, develop, and optimize PySpark-based ETL pipelines running on onprem Hadoop clusters and cloud environments.
  • Build highvolume ingestion frameworks using Kafka for real-time and near-real-time trading and market data.
  • Develop, tune, and manage Hadoop ecosystem components HDFS, YARN, MapReduce, Tez, Oozie/Airflow.
  • Build high-performance, optimized Hive data models for regulatory reporting, trade lifecycle, and market risk processing.

Databricks Lakehouse & Delta Framework

  • Architect and implement Bronze/Silver/Gold layer modeling patterns within the Databricks Lakehouse.
  • Apply Delta Lake best practices including:

o optimized file management

o Z-Ordering

o Delta Change Data Feed (CDF)

o schema evolution & enforcement

o ACID transaction handling

  • Build reusable frameworks for ingestion, cleansing, transformation, and consumption of data across Lakehouse layers.
  • Enable governance, lineage, and auditability using Unity Catalog or equivalent cataloging tools.

Collaboration, Leadership & Delivery

  • Collaborate closely with quants, product owners, architects, risk tech, and business users.
  • Participate in agile ceremonies sprint planning, refinement, design reviews.
  • Mentor junior engineers and contribute to building strong engineering practices across tech teams.

Requirements

Must Have Technical/Functional Skills

Primary skills: Hadoop, Hive, Python, PySpark, Apache Kafka, Hadoop Ecosystem, Hive, Databricks Lakehouse Architecture, Delta Lake, Bronze/Silver/Gold Data Modeling, Big Data ETL Pipeline Development, SQL, Real-time Data Ingestion Frameworks, Data Governance & Cataloging, CI/CD Tools Git, Jenkins, Bitbucket, Workflow Orchestration, and Cloud & On-Prem Big Data Platforms.

Roles & Responsibilities

Seeking a Senior Big Data Engineer with 9-14 years of experience specializing in Hadoop, Python, Hive PySpark, Kafka, and strong experience designing data solutions for large-scale financial systems.

In addition, the candidate must possess advanced expertise in Databricks Lakehouse architecture, particularly around Bronze/Silver/Gold layer data modeling, Delta Lake optimizations, and building reliable, scalable pipelines for regulatory, risk, trading, and analytics workloads., * 9-14 years of hands-on experience in Big Data Engineering.

  • Expert skills in:

o PySpark data frame optimizations, partitioning, broadcast strategies, distributed computing.

o Kafka producer/consumer design, schema registry, streaming ETLs.

o Hadoop ecosystem HDFS, YARN, MapReduce/Tez, Oozie/Airflow.

o Hive advanced query tuning, TEZ optimization, partition/bucket management.

  • Extensive hands-on experience with Databricks Lakehouse, including:

o Bronze/Silver/Gold layer modeling

o Delta Lake optimizations

o Data quality frameworks on Lakehouse

o Structured & unstructured data handling

  • Experience in Global Markets, Risk, Treasury, Trade Surveillance, or Regulatory Reporting.
  • Strong SQL knowledge with experience working on massive datasets (TB/PB scale).
  • Experience with CI/CD practices Git, Jenkins, Bitbucket, build pipelines.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

1:02 min

Applying an ETL methodology to infrastructure configuration management

Axel Barbier · World Congress 2023

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

57 sec

Extracting API schemas automatically during continuous integration builds

Axel Barbier · World Congress 2023

Videos

See all

Related articles

See all