Senior Data Engineer - Azure / Fraud Analytics

Info Gain Consulting LLC
United States
17 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Compensation
$123,533.0 - $148,771.0
Working hours
Regular working hours
Job source

Tech stack

Data Analysis Microsoft Azure Cloud Database Continuous Integration Data Architecture Data Dictionary Extract Transform Load (ETL) Data Warehousing Python (Programming Language) Machine Learning Modular Design Azure Machine Learning
+13 more
Azure Data Lake SQL Databases Transact-SQL Parquet Data Logging Data Processing Large Language Models Pandas Pyspark Machine Learning Operations Restful APIs Azure Synapse Analytics Software Version Control

Job description

We are hiring a Senior Data Engineer to build and maintain the Azure-based data architecture that powers fraud analytics for a federal Office of Inspector General (OIG). You will design integrated, flexible, and secure pipelines that move source data into a modern cloud environment and make it available for audits, investigations, and machine learning.

You’ll work code-first - source control, CI/CD, and reusable modular design are core to how this team operates. If you enjoy owning architecture end to end and partnering with data scientists to make ML pipelines fast and reliable, this role is for you.

What You’ll Do

  • Design, implement, and maintain an efficient, secure, stable, and flexible data architecture, with all assets managed via source control
  • Migrate source data to Azure Data Lake Storage (ADLS) and build/maintain ELT/ETL pipelines in Azure Synapse and Azure Machine Learning (SDK V1 and V2)
  • Review and improve existing architecture and pipelines - periodic audits to address bottlenecks, deprecated dependencies, and architecture drift
  • Establish quality controls, error handling, logging, and validation checks across all pipelines
  • Normalize entity attributes (addresses, phone numbers, and other common fields)
  • Optimize ingestion, processing, and storage across diverse datasets, including modern columnar formats such as Parquet
  • Build self-service capabilities so analysts can query and export data for investigations and audits
  • Partner with data scientists to ensure the architecture efficiently supports machine learning workloads
  • Author and maintain SOPs governing authoring, validation, publishing, execution, and monitoring of all pipelines and assets
  • Produce detailed documentation - data dictionaries, ER diagrams, and pipeline process maps
  • Stay current with emerging AI tooling and contribute to efforts evaluating automation and LLM-assisted capabilities, * Core hours between 6:00 a.m. and 6:00 p.m. local time, Monday-Friday
  • Telework authorized; occasional on-site meetings and collaboration as requested
  • Government-furnished equipment provided for work on government systems
  • Federal holidays observed

Pay: $123,532.67 - $148,770.53 per year

Requirements

  • 5 years of hands-on experience in each of:

  • Maintaining SQL databases and conducting advanced operations in SQL and T-SQL

  • Designing, implementing, and maintaining ELT/ETL processes in cloud-based data analytics environments
  • 3 years of hands-on experience in each of:

  • Working in Azure Synapse and Azure Machine Learning with the modern data stack - certifications preferred (DP-203 or equivalent)

  • Manipulating data in Python (Pandas required; PySpark/Polars preferred; experience with reusable, modular code preferred)

Preferred

  • Implementing pipelines and infrastructure using code-first approaches (Python SDK, CLI, REST APIs, or IaC tooling)
  • Implementing source control and CI/CD workflows
  • Demonstrated familiarity with AI coding assistants and LLM integration patterns
  • Azure certification (DP-203 or equivalent), * Maintaining SQL databases: 5 years (Preferred)
  • Conducting advanced operations in SQL and T-SQL: 5 years (Preferred)
  • Design & maintaining ELT/ETL processes : 5 years (Preferred)
    • Manipulating data in Python (Pandas required): 3 years (Preferred)

License/Certification:

  • Azure certification DP-203 or equivalent (Required)

Benefits & conditions

3.73.7 out of 5 stars Remote $123,532.67 - $148,770.53 a year - Full-time, Pulled from the full job description

  • 401(k)
  • Health insurance
  • 401(k) matching
  • Vision insurance
  • Dental insurance
  • Life insurance, * 401(k)
  • 401(k) matching
  • Dental insurance
  • Health insurance
  • Life insurance
  • Vision insurance

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:50 min

How Parquet metadata enables efficient data reading

Matthias Niehoff Matthias Niehoff · WWC Europe 2026

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

2:03 min

Accelerating pandas dataframes using cudf module plugins

Ankit Patel Ankit Patel · WWC 2024

3:33 min

Refactoring data science workflows using Rapids QDF and Pandas

Paul Graham Paul Graham · LIVE

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

Videos

See all

Related articles

See all