Streaming Data

OpenKyber LLC
United States
3 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Airflow Amazon Web Services Business Analytics Applications Apache HTTP Server Microsoft Azure Big Data Cloud Computing Computer Programming Data Architecture Information Engineering Data Governance
+22 more
Data Systems Apache Hive Python (Programming Language) Machine Learning Performance Tuning Query Optimization Cloud Services SQL Databases Data Streaming Workflow Management Systems Google Cloud Snowflake Apache Spark Indexer Data Lakes Infrastructure Automation Frameworks Collibra Apache Kafka Spark Streaming Data Lakehouse Data Pipelines Databricks

Job description

We are seeking a highly skilled Senior Lead Data Engineer with strong experience in modern data platforms including Snowflake , Databricks , Apache Iceberg , and Apache Spark . The ideal candidate will lead the design, development, and optimization of scalable data pipelines and analytics platforms while ensuring high performance for large-scale SQL workloads . This role requires strong expertise in data architecture, performance tuning, and big data technologies to support enterprise-level analytics and data-driven decision-making., * Design and implement scalable data pipelines and data lakehouse architectures using Snowflake, Databricks, and Apache Iceberg.

  • Lead the development and optimization of Spark-based ETL/ELT pipelines for large-scale data processing.
  • Optimize complex SQL workloads for performance, cost efficiency, and scalability.
  • Build and maintain high-performance data models supporting analytics, reporting, and machine learning workloads.
  • Implement data governance, security, and data quality frameworks.
  • Collaborate with data scientists, analysts, and business stakeholders to deliver reliable data solutions.
  • Perform performance tuning for distributed processing frameworks such as Spark and Databricks.
  • Guide engineering teams on best practices for data architecture, pipeline orchestration, and cloud data platforms .
  • Monitor and troubleshoot data pipeline performance and reliability issues.
  • Mentor junior data engineers and lead technical design discussions.

Requirements

  • 10+ years of experience in Data Engineering or Big Data Engineering .
  • Strong expertise with Snowflake and Databricks Lakehouse platform .
  • Hands-on experience with Apache Spark (PySpark / Spark SQL) .
  • Experience working with Apache Iceberg or modern table formats .
  • Advanced knowledge of SQL performance tuning and query optimization .
  • Experience designing data lake / lakehouse architectures .
  • Strong programming experience in Python, Scala, or Java .
  • Experience with workflow orchestration tools (Airflow, Prefect, or similar).
  • Knowledge of cloud platforms such as Amazon Web Services , Microsoft Azure , or Google Cloud .
  • Strong understanding of data modeling, partitioning, indexing, and storage optimization.

Preferred Qualifications

  • Experience with data lakehouse architecture and open table formats .
  • Knowledge of streaming data pipelines using Kafka or Spark Streaming.
  • Experience with CI/CD pipelines and infrastructure-as-code tools .
  • Strong leadership and mentoring experience.
  • Experience supporting enterprise-scale analytics platforms .

Nice to Have

  • Experience with data governance tools.
  • Knowledge of machine learning data pipelines.
  • Certifications in cloud platforms or data engineering technologies.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · WWC Europe 2026

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:36 min

Analyzing limitations with PostgreSQL bitmap heap scans

Dharin Shah Dharin Shah · WWC 2025

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · WWC Europe 2026

3:05 min

Audience questions on AI agents and pipeline vectorization

Joy Joy · WWC 2024

Videos

See all

Related articles

See all