Lead AI Data Engineer

Insight Global
Frisco, TX, United States
about 2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$200,000.0 - $220,000.0
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services Amazon S3 Business Analytics Applications Cyber Security Extract Transform Load (ETL) Database Queries Distributed Systems Apache Hive Python (Programming Language) Operational Databases
+14 more
Performance Tuning Power BI SQL Databases Tableau (Software) Enterprise Software Applications Grafana Apache Spark Data Lakes Pyspark Data Analytics Operational Systems Data Pipelines Databricks Microservices

Job description

Insight Global is seeking a Lead AI Data Engineer to sit hybrid at a Cybersecurity client in Frisco, Texas. You will join their eCommerce Operational Intelligence team, building enterprise-scale data pipelines and analytics foundations (SQL, Spark/PySpark, ETL/ELT) that produce reliable operational insights and measurable business impact.

You will drive the transformation of eCommerce operational analytics and real-time monitoring by building scalable data pipelines, AI-powered insights, and intelligent dashboards. This role leads AI proof-of-concepts and contributes to production-grade solutions that improve platform reliability, accelerate root-cause identification, enhance engineering productivity, and strengthen operational intelligence across their eCommerce ecosystem.

Day to Day:

-Build and operate production ETL/ELT pipelines processing millions of eCommerce events daily and order trends.

-Write and tune complex SQL for operational analytics, KPIs, and reporting.

-Design analytics-ready schemas and data models for performance and scale.

-Troubleshoot pipelines, microservices, and APIs; apply observability to isolate root causes.

-Integrate data across eCommerce, MarTech, and operational systems into unified insights.

Requirements

10+ years building and architecting large-scale applications and distributed systems.

-5+ years building production data pipelines, ETL/ELT workflows, and analytics platforms.

-Applied AI to operational intelligence (anomaly detection/alerting, forecasting, insights).

-Expert SQL (complex queries, performance tuning)

-Spark/PySpark in production (Spark SQL, optimization)

-Strong Python (testing, packaging, best practices)

-ETL/ELT pipelines (orchestration, monitoring, error handling)

-Databricks & Delta Lake (Jobs, Unity Catalog, Medallion)

-Analytics data modeling (star/snowflake schemas)

-Distributed systems (APIs, microservices, event-driven)

-Production observability & troubleshooting

-AWS (S3, Lambda, Glue, Kinesis, OpenSearch, QuickSight), BI (Power BI, Tableau, Grafana), GenAI (RAG, vector DBs, LangChain, Bedrock)

Benefits & conditions

(This role can pay $200,000 - 220,000 based on years of experience).

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on juju.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

1:46 min

Traditional data architecture before Microsoft Fabric

Dr. Alexander Wachtel Dr. Alexander Wachtel +1 · WWC 2025

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

1:43 min

AWS infrastructure stack and data flow pipeline overview

Artem Volk Artem Volk +1 · WWC 2024

Videos

See all

Related articles

See all