ETL/Data Engineer

VIRGENCE GROUP LLC
Indianapolis, IN, United States
26 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
2 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Data Analysis Unit Testing Microsoft Azure Bash Shell Code Review Continuous Integration Data Systems Data Vault Modeling Data Warehousing DevOps Apache Hive
+44 more
Key Management Log Analysis Microsoft SQL Server SQL Azure Oracle (Applications) Performance Tuning Windows PowerShell Query Optimization Role-Based Access Control Release Management Azure Active Directory Kusto Query Language Azure Data Lake SQL Databases SQL Server Integration Services Data Streaming Teradata SQL Transact-SQL YAML Parquet Data Logging Scripting Azure Data Factory Sql Optimization Cloud Monitoring Informatica Powercenter Snowflake Apache Spark Change Data Capture Pandas Build Management Microsoft Fabric Pytest Data Lakes Pyspark Git Flow Bicep Software Coding Restful APIs Terraform Azure Synapse Analytics Data Pipelines Serverless Computing Databricks

Job description

enterprise data platform on Microsoft Azure. You will own end-to-end delivery of data pipelines and data products that power analytics, regulatory reporting, operational dashboards, and emerging AI/ML use cases. You will partner closely with data architects, analytics engineers, data scientists, business stakeholders, and platform engineering teams to deliver reliable, performance, secure, and costefficient data solutions. This role is ideal for an engineer with strong hands-on depth in Azure Data Factory, Azure Synapse Analytics and/or Databricks, and modern Lakehouse patterns, who is comfortable leading migration programs (e.g., Informatica-to-ADF, on-prem warehouse-to-cloud), mentoring mid-level engineers, and shaping engineering standards across the team., Pipeline Design & Development Design and build robust, reusable, parameter-driven ingestion and transformation pipelines using Azure Data Factory, Synapse Pipelines, Data Bricks and/or Microsoft Fabric Data Factory. Implement medallion architecture (Bronze / Silver / Gold) on Azure Data Lake Storage Gen2 using Delta Lake, Parquet, and structured streaming patterns. Build performant ELT workflows that leverage pushdown to source systems (Synapse Dedicated SQL Pool, Azure SQL, Teradata) where appropriate. Develop and optimize PySpark notebooks and jobs on Azure Databricks or Synapse Spark. Data Modeling & Warehousing Design dimensional models (Kimball star/snowflake) and data vault patterns for analytics consumption. Implement Slowly Changing Dimensions (Type 1/2/3), Change Data Capture, and late-arriving data patterns. Tune distributed SQL workloads in Synapse Dedicated SQL Pool / Fabric Warehouse, including distribution keys, partitioning, and clustered column store indexes. Platform Engineering & DevOps Implement CI/CD for data pipelines using Azure DevOps (YAML pipelines, ARM/Bicep/Terraform) across Dev / SIT / UAT / Prod environments. Instrument pipelines with robust logging, auditing, and monitoring using Azure Monitor, Log Analytics, and KQL. Define and enforce coding standards, code review practices, branching strategies, and release management. Migration & Modernization Lead or contribute to legacy-to-cloud migrations - e.g., Informatica PowerCenter to Azure Data Factory, on-premises Teradata / Oracle / SQL Server to Synapse or Fabric. Perform workload assessment, capacity planning, and cost modeling for target-state architectures. production incident response for critical pipelines.

Requirements

Deep hands-on expertise with Azure Data Factory: pipelines, datasets, linked services, triggers, parameterization, mapping data flows, and all three Integration Runtime types (Azure, Selfhosted, SSIS). Strong Experience in Data Bricks and PySpark. Production experience with one or more of: Azure Synapse Analytics (Dedicated and Serverless SQL Pools, Spark Pools) OR Azure Databricks (Delta Lake, Unity Catalog) OR Microsoft Fabric (Warehouse, Lakehouse, OneLake). Strong working knowledge of Azure Data Lake Storage Gen2 (hierarchical namespace, RBAC + ACLs, lifecycle management, security). Experience with Azure Key Vault, Azure AD / Entra ID (including managed identities and service principals), and private networking (VNet integration, private endpoints). Monitoring and troubleshooting with Azure Monitor, Log Analytics, and KQL. Advanced SQL - window functions, CTEs, query optimization, execution plan analysis, performance tuning. Strong Python for data engineering - pandas, PySpark, REST API integration, unit testing (pytest). Proficient in T-SQL; familiarity with Spark SQL, KQL, PowerShell, and Bash shell scripting.

Required Qualifications: Deep hands-on expertise with Azure Data Factory: pipelines, datasets, linked services, triggers, parameterization, mapping data flows, and all three Integration Runtime types (Azure, Selfhosted, SSIS). Strong Experience in Data Bricks and PySpark. Production experience with one or more of: Azure Synapse Analytics (Dedicated and Serverless SQL Pools, Spark Pools) OR Azure Databricks (Delta Lake, Unity Catalog) OR Microsoft Fabric (Warehouse, Lakehouse, OneLake). Strong working knowledge of Azure Data Lake Storage Gen2 (hierarchical namespace, RBAC + ACLs, lifecycle management, security). Experience with Azure Key Vault, Azure AD / Entra ID (including managed identities and service principals), and private networking (VNet integration, private endpoints). Monitoring and troubleshooting with Azure Monitor, Log Analytics, and KQL. Advanced SQL - window functions, CTEs, query optimization, execution plan analysis, performance tuning. Strong Python for data engineering - pandas, PySpark, REST API integration, unit testing (pytest). Proficient in T-SQL; familiarity with Spark SQL, KQL, PowerShell, and Bash shell scripting. Preferred Qualifications: 5+ years of data warehouse development experience. 5+ years of data modeling experience using ERWIN or similar tools. 2+ years of experience with Azure Data Factory and Snowflake. Medicaid Domain Knowledge is a plus

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

1:35 min

Centralizing configuration logic with native YAML block references

Matthieu Vincent Matthieu Vincent · Europe 2026 Virtual

2:56 min

Provisioning a secure container infrastructure with Bicep

Matthias Falkenberg +1 · WWC 2022

3:05 min

Tagging and organizing execution scenarios with pytest markers

Florian Bruhin · WWC 2021

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · WWC Europe 2026

2:00 min

Introduction to YAML syntax and basic formatting

Chris Ayers · LIVE

Videos

See all

Related articles

See all