Data/Scala/Spark Engineering Specialist

Anagha Techno Soft
New York, NY, United States
2 months ago

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Airflow Unit Testing CA Workload Automation Ae Microsoft Azure Continuous Integration Information Engineering Extract Transform Load (ETL) Github Python (Programming Language) Networking Basics Performance Tuning
+9 more
Cloud Services SQL Databases Teradata SQL Snowflake Apache Spark Star Schema Serverless Computing Databricks Artifactory

Job description

We’re migrating complex on-prem regulatory reporting pipelines from a legacy ETL + Autosys + SQL + Teradata stack to a modern Databricks + Snowflake platform on Azure. The role is hands-on: design, implement, test, and reconcile production pipelines feeding regulatory reports under strict parity requirements.

Requirements

Scala / Spark production experience writing Spark applications in Scala (not just notebooks); comfortable with the Data Frame API, joins, window functions, partitioning, and performance tuning Databricks Serverless compute, Unity Catalog, Asset Bundles, Databricks CLI SQL fluency comfortable writing, analyzing and extracting requirements from complex SQL scripts Snowflake schema design, performance, Spark-Snowflake connector Azure ADLS, networking basics, secrets/identity (Entra ID / managed identities) Orchestration Airflow (DAG authoring, sensors, retries, SLAs) CI/CD Artifactory, GitHub Actions pipelines: build, sharded test matrices, artifact promotion through dev QA UAT prod Testing Experience in TDD, writing unit tests (ScalaTest, AnyFlatSpec) and BDD (Concordion or equivalent) Data quality & reconciliation building automated parity checks against legacy outputs, drift detection, row-level reconciliation tooling Large-scale migrations proven track record migrating legacy ETL (Autosys/Informatica/etc.) to cloud data platforms, including dependency mapping and cutover planning Modern data engineering practices medallion architecture (Bronze/Silver/Gold), idempotent pipelines, schema evolution, lineage, observability

Nice-to-have

Financial services / regulatory reporting domain Python (Databricks utilities, tooling) Spec-driven development workflows (specs plans tasks implementation)

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · WWC Europe 2026

6:08 min

Applying software engineering environments and testing to data pipelines

Matthias Niehoff Matthias Niehoff · WWC 2024

1:33 min

Integrating internal APIs and maintaining data sovereignty

Mahran Meißner Mahran Meißner · WWC Europe 2026

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · WWC 2023

2:34 min

Capabilities of the Apache Spark processing engine

Ayon Roy · LIVE

2:46 min

Transforming data architecture from on-premise to cloud

Sandhya Menon Sandhya Menon · WWC Europe 2026

Videos

See all

Related articles

See all