Lead BI Engineer

CRG
United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Microsoft Azure Big Data Continuous Integration Data Architecture Information Engineering Data Governance Data Security Distributed Data Store Python (Programming Language) Machine Learning Performance Tuning
+13 more
Cloud Services DataOps Azure Machine Learning SQL Databases Data Streaming Azure Service Bus Apache Spark Data Strategy Data Lakes Apache Kafka Machine Learning Operations Tools for Reporting Data Pipelines

Job description

At CRG, we are seeking an experienced Lead BI Engineer to lead the design, development, and optimization of scalable Business Intelligence solutions. This role is responsible for driving data strategy, building robust data models and reporting platforms, and mentoring BI engineers while partnering with cross-functional stakeholders to deliver actionable insights that support business decision-making., * Architect, build, and optimize distributed data pipelines using Apache Spark in a high-volume, mission-critical environment.

  • Design and maintain enterprise Lakehouse architecture with Delta Lake, ensuring ACID compliance, lineage, auditability, and data governance.
  • Develop automated ingestion frameworks (batch, streaming, and event-driven) across multiple cloud services and integration points.
  • Enable machine-learning workflows by preparing feature-ready datasets and establishing reproducible ML deployment patterns.
  • Lead platform-wide data quality, access control, and cataloging frameworks.
  • Implement advanced cost-optimization, cluster tuning, and performance engineering strategies.
  • Collaborate with Finance, BI, Operations, and ML teams to translate complex business needs into scalable data solutions.
  • Own production reliability, troubleshooting, and root-cause analysis for data and ML pipelines.

Requirements

  • 7+ years of experience in advanced data engineering with distributed compute technologies.
  • Expert-level Spark engineering (performance tuning, cluster configuration, partition strategies, optimization of large datasets).
  • Hands-on experience with Lakehouse architectures including ACID transactions, schema evolution, and governance frameworks.
  • Deep proficiency in Python and SQL for large-scale data transformation.
  • Experience supporting machine-learning pipelines or model operationalization.
  • Proven experience architecting cloud-native data platforms (Azure, AWS, or GCP).
  • Strong background integrating diverse, complex data sources at enterprise scale.
  • Demonstrated ability to own mission-critical production systems.
  • Experience with distributed streaming frameworks (Kafka, Event Hubs, or similar).
  • Experience building or supporting ML platforms, feature stores, or experiment-tracking systems.
  • Background in data security, compliance controls, or audit-ready governance.
  • Experience automating data operations with CI/CD and infrastructure-as-code.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:24 min

The governance failures of centralized data lakes

Mario Meir-Huber · LIVE

2:00 min

Separating dataset creation from low-level software implementation steps

Jan Zawadzki · WWC 2022

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

6:24 min

Distributed data lakes and containerized computing clusters

Ulrich Wurstbauer +1 · LIVE

2:56 min

Core data lake environment and infrastructure requirements

Christoph Fassbach Christoph Fassbach +1 · WWC 2024

Videos

See all

Related articles

See all