Lead BI Engineer
CRG
United States
about 1 month ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours
Job source
Tech stack
Amazon Web Services
Microsoft Azure
Big Data
Continuous Integration
Data Architecture
Information Engineering
Data Governance
Data Security
Distributed Data Store
Python (Programming Language)
Machine Learning
Performance Tuning
+13 more
Cloud Services
DataOps
Azure Machine Learning
SQL Databases
Data Streaming
Azure Service Bus
Apache Spark
Data Strategy
Data Lakes
Apache Kafka
Machine Learning Operations
Tools for Reporting
Data Pipelines
Job description
At CRG, we are seeking an experienced Lead BI Engineer to lead the design, development, and optimization of scalable Business Intelligence solutions. This role is responsible for driving data strategy, building robust data models and reporting platforms, and mentoring BI engineers while partnering with cross-functional stakeholders to deliver actionable insights that support business decision-making., * Architect, build, and optimize distributed data pipelines using Apache Spark in a high-volume, mission-critical environment.
- Design and maintain enterprise Lakehouse architecture with Delta Lake, ensuring ACID compliance, lineage, auditability, and data governance.
- Develop automated ingestion frameworks (batch, streaming, and event-driven) across multiple cloud services and integration points.
- Enable machine-learning workflows by preparing feature-ready datasets and establishing reproducible ML deployment patterns.
- Lead platform-wide data quality, access control, and cataloging frameworks.
- Implement advanced cost-optimization, cluster tuning, and performance engineering strategies.
- Collaborate with Finance, BI, Operations, and ML teams to translate complex business needs into scalable data solutions.
- Own production reliability, troubleshooting, and root-cause analysis for data and ML pipelines.
Requirements
- 7+ years of experience in advanced data engineering with distributed compute technologies.
- Expert-level Spark engineering (performance tuning, cluster configuration, partition strategies, optimization of large datasets).
- Hands-on experience with Lakehouse architectures including ACID transactions, schema evolution, and governance frameworks.
- Deep proficiency in Python and SQL for large-scale data transformation.
- Experience supporting machine-learning pipelines or model operationalization.
- Proven experience architecting cloud-native data platforms (Azure, AWS, or GCP).
- Strong background integrating diverse, complex data sources at enterprise scale.
- Demonstrated ability to own mission-critical production systems.
- Experience with distributed streaming frameworks (Kafka, Event Hubs, or similar).
- Experience building or supporting ML platforms, feature stores, or experiment-tracking systems.
- Background in data security, compliance controls, or audit-ready governance.
- Experience automating data operations with CI/CD and infrastructure-as-code.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.indeed.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
DS
Dhannush Subramani
about 4 years ago
BB
Benedikt Bischof
Making Data Warehouses Fast: A Developer’s Story
about 4 years ago
EM
Eli McGarvie
Data Engineer Salary UK
about 3 years ago
CH
Chris Heilmann
Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production
almost 2 years ago
EM
Eli McGarvie
Highest Paying Tech Companies for Developers
over 3 years ago
MH
Michael Hunger
Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?
7 months ago