Data Platform Engineer

Bright Vision Technologies
Durham, NC, United States
13 days ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$100,000.0 - $150,000.0
Working hours
Regular working hours

Tech stack

Java (Programming Language) Airflow Apache HTTP Server Microsoft Azure Big Data Cloud Computing Continuous Integration Information Engineering Data Governance Data Infrastructure Data Stores Database Queries
+25 more
Software Debugging Distributed Systems Fault Tolerance Apache Hadoop Hadoop Distributed File System Apache HBase Apache Hive Python (Programming Language) Machine Learning NoSQL Apache Oozie Sqoop Data Streaming Unstructured Data Scripting Apache Spark Hdinsight Kubernetes Collibra Apache Flink Apache Kafka Spark Streaming Data Management Amazon Elastic Mapreduce (EMR) Databricks

Job description

We are seeking an experienced Data Platform Engineer to design, build, and operate large-scale data processing pipelines and analytics platforms on Hadoop and related big-data ecosystems. In this role you will be responsible for ingesting, transforming, and analyzing massive volumes of structured and unstructured data to support enterprise analytics, machine learning, and reporting workloads. The ideal candidate will combine deep technical expertise across the Hadoop ecosystem with strong software engineering fundamentals and a clear understanding of how to deliver reliable, performant, and cost-effective data platforms in production environments.

Requirements

  • 5+ years of professional experience designing and operating big-data pipelines on Hadoop.
  • Strong hands-on expertise with Apache Spark (Scala, Python, or Java) in production environments.
  • Solid experience with Hive, HDFS, Sqoop, HBase, and the broader Hadoop ecosystem.
  • Hands-on experience with streaming data platforms such as Kafka, Spark Streaming, or Flink.
  • Strong SQL skills and experience working with both relational and NoSQL data stores.
  • Experience with workflow orchestration tools such as Airflow or Oozie.
  • Solid understanding of distributed systems concepts, including partitioning, replication, and fault tolerance.
  • Strong scripting skills in Python or Shell.
  • Excellent troubleshooting, debugging, and documentation skills.

Preferred Qualifications

  • Experience operating Hadoop on cloud platforms such as AWS EMR, Azure HDInsight, or Databricks.
  • Familiarity with modern lakehouse formats (Delta, Iceberg, Hudi).
  • Exposure to data governance tooling such as Apache Atlas or Collibra.
  • Experience with Kubernetes-based data platforms (Spark-on-K8s, Trino).
  • Hands-on experience with CI/CD and infrastructure-as-code in data engineering workflows.

Benefits & conditions

  • $146,200-189,200 per year At Gilead, we’re creating a healthier world for all people. For more than 35 years, we’ve tackled diseases such as HIV, viral hepatitis, COVID-19 and cancer - working relentlessly …

  • 27 days ago +

About the company

Gilead

  • Raleigh, NC
  • $146,200-189,200 per year At Gilead, we’re creating a healthier world for all people. For more than 35 years, we’ve tackled diseases such as HIV, viral hepatitis, COVID-19 and cancer - working relentlessly …

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:37 min

Comparing traditional SQL tables versus NoSQL non-tabular databases

Stanimira Vlaeva · JS Congress

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

3:16 min

Terminology differences between relational and NoSQL databases

Tim Faulkes · LIVE

3:05 min

Audience questions on AI agents and pipeline vectorization

Joy Joy · World Congress 2024

Videos

See all

Related articles

See all