Lead Data Engineer

Raas Infotek LLC
Manor, United States
6 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Airflow Amazon Web Services Amazon S3 Microsoft Azure Big Data BigQuery Information Systems Computer Programming Continuous Integration Data as a Services Information Engineering
+43 more
Data Governance Data Infrastructure Extract Transform Load (ETL) Data Vault Modeling Data Warehousing Software Debugging Data Flow Control Apache Hadoop Apache Hive Python (Programming Language) Operational Databases Performance Tuning Query Optimization Cloud Services DataOps Cloudera SQL Databases Data Streaming Unstructured Data Google Cloud Azure Data Factory Snowflake Data Build Tool (dbt) Apache Spark Git Cloudformation Containerization Data Lakes Kubernetes Information Technology Apache Flink Data Analytics Star Schema Apache Kafka Spark Streaming Data Management Machine Learning Operations Api Design Terraform Software Version Control Data Pipelines Docker Databricks

Job description

We are looking for a highly experienced Data Engineer to design, build, and optimize scalable data platforms and pipelines. The ideal candidate has deep expertise across the modern data stack, strong architectural judgment, and the ability to lead data engineering initiatives end-to-end while mentoring junior engineers and collaborating closely with data science, analytics, and business teams., * Design, build, and maintain scalable, reliable ETL/ELT pipelines for structured and unstructured data

  • Architect and optimize data warehouses, data lakes, and lakehouse solutions
  • Lead the design of data models (dimensional, normalized, and denormalized) to support analytics and reporting
  • Build and manage batch and real-time streaming data pipelines
  • Ensure data quality, integrity, governance, and security across all pipelines and platforms
  • Optimize pipeline performance, cost, and scalability across cloud environments
  • Collaborate with data scientists, analysts, and business stakeholders to understand data requirements
  • Implement CI/CD practices for data pipelines and infrastructure-as-code
  • Monitor, troubleshoot, and resolve production data pipeline issues (on-call/incident support as needed)
  • Establish and enforce data engineering best practices, standards, and documentation
  • Lead technical design reviews and mentor junior/mid-level data engineers
  • Evaluate and integrate new tools/technologies to improve the data platform

Requirements

  • 10+ years of experience in data engineering, data warehousing, or related fields
  • Strong programming skills in Python, Scala, or Java
  • Expert-level SQL and experience with query optimization on large datasets
  • Hands-on experience with big data technologies: Spark, Hadoop, Kafka, Hive
  • Strong experience with cloud data platforms - AWS (Redshift, Glue, EMR, S3), Azure (Synapse, Data Factory, Databricks), or Google Cloud Platform (BigQuery, Dataflow, Dataproc)
  • Experience with modern data warehouse/lakehouse platforms - Snowflake, Databricks, or similar
  • Proficiency in orchestration tools - Airflow, Dagster, or similar
  • Solid understanding of data modeling (Star/Snowflake schema, Data Vault) and dimensional design
  • Experience with real-time/streaming architectures (Kafka, Kinesis, Flink, Spark Streaming)
  • Strong knowledge of CI/CD, version control (Git), and infrastructure-as-code (Terraform/CloudFormation)
  • Experience with containerization and orchestration (Docker, Kubernetes)
  • Understanding of data governance, security, lineage, and compliance (GDPR/HIPAA as applicable)
  • Proven experience architecting solutions from scratch and leading data engineering teams/projects
  • Excellent problem-solving, debugging, and performance-tuning skills

Good to Have

  • Experience with dbt (data build tool) for transformation workflows
  • Exposure to MLOps and supporting ML pipelines/feature stores
  • Experience with DataOps practices and data observability tools (Monte Carlo, Great Expectations)
  • Knowledge of API development for data services
  • Relevant cloud certifications (AWS Certified Data Analytics, Azure Data Engineer Associate, Google Cloud Platform Professional Data Engineer)
  • Experience in a specific domain (Finance, Healthcare, Retail, etc. - customize as needed)

Educational Qualification

  • Bachelor’’s/Master’’s degree in Computer Science, Data Engineering, Information Systems, or related field, * Strong communication skills to work with cross-functional and non-technical stakeholders
  • Leadership ability to mentor and guide junior engineers
  • Ability to drive projects independently with minimal supervision
  • Strong ownership mindset and attention to data quality/detail

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:05 min

Audience questions on AI agents and pipeline vectorization

Joy Joy · World Congress 2024

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

Videos

See all

Related articles

See all