Data Engineer

Inclusion Inc.
Dallas, TX, United States
3 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Airflow Microsoft Azure Software as a Service Databases Extract Transform Load (ETL) Data Warehousing Programming Tools Revision Control Systems Python (Programming Language) Machine Learning Power BI
+16 more
SQL Databases Transact-SQL Workflow Management Systems Datadog Apache Spark Git Microsoft Fabric Pyspark Performance Monitor Database Mirroring Cloud Optimization Restful APIs Software Version Control Data Pipelines Custom Reports Databricks

Job description

Our client is building a next generation SaaS application that ingests data from various sources, processes through a data pipeline and executes machine learning models to provide predictions on Project Management data. The solution will leverage cutting edge technologies hosted in Microsoft Azure including Databricks, Apache Airflow, Fabric database mirroring, Fabric Lakehouse, Semantic Models, Materialized Lake Views and Power BI. The app users will have the ability to access Machine Learning predictions, create custom reports and combine data from various sources.

Requirements

Do you have experience in Version control?, * Python (expert)

  • Spark/pySpark (advanced)
  • T-SQL (advanced)
  • ETL/ELT and creation of data pipelines (expert)
  • Databricks (advanced)
  • Data modeling (advanced)
  • Experience with scheduling and orchestration tools, specifically Apache Airflow (advanced)
  • Azure services including compute, storage, databases and developer tools (advanced)
  • Data warehousing concepts (expert)
  • Version control tools such as Azure DevOps Git (advanced)
  • Performance monitoring and optimization of code and tooling (advanced)
  • Security including row level and object level
  • Invoking RESTful API’s
  • Strong aptitude and ability to work independently
  • Medallion architecture

Desired Skills

  • Experience with Microsoft Fabric and implementing data pipelines using Fabric tooling
  • Knowledge of AI concepts and experience executing ML models
  • Cloud cost optimization
  • Knowledge of OneLake security
  • Sharing of Power BI and Semantic Models in Microsoft Fabric
  • Understanding of cross tenant data sharing in Microsoft Fabric
  • Team leadership experience
  • Observability tooling, * Knowledge of Project Management and Financial concepts including budgets, tasks, revenue, profit and earned value
  • Certification in Microsoft Azure and/or Python, SQL, Databricks

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · WWC Europe 2026

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

3:05 min

Audience questions on AI agents and pipeline vectorization

Joy Joy · WWC 2024

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

Videos

See all

Related articles

See all