AWS Databricks Data Engineer

Xoriant Corporation
United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Airflow Amazon Web Services Amazon Elastic Compute Cloud Amazon S3 Apache HTTP Server CA Workload Automation Ae Cloud Computing Cloud Database Program Optimization Continuous Integration Data Architecture
+40 more
Data Governance Data Integrity Extract Transform Load (ETL) Data Visualization Data Warehousing Relational Databases Software Debugging DevOps Distributed Data Store Apache Hive Identity and Access Management Python (Programming Language) Machine Learning Automation of Marketing Meta-Data Management Performance Tuning Power BI Scala (Programming Language) Scripting Data Classification Data Ingestion Microsoft Power Automate Sql Optimization Apache Spark Caching Indexer Git Data Lakes Pyspark Data Lineage Deployment Automation Real Time Data Cloudwatch Software Coding Software Version Control Data Pipelines Serverless Computing Powerapps Docker Databricks

Job description

We are seeking a highly skilled Cloud Data Engineer to design, build, and optimize a modern, scalable Legal Data Lakehouse platform. Operating within State Street’s Global Technology Services, you will leverage a deep knowledge of the full suite of AWS cloud services combined with high-performance Databricks capabilities to ingest, model, and secure complex enterprise data structures (including contracts, litigation matters, eDiscovery datasets, and global regulatory feeds).

This role is critical to establishing a single, highly governed, audit-ready source of truth that powers critical legal operations, compliance analytics, and emerging generative AI/ML use cases across our global footprint., * Design, build, and maintain enterprise-grade, custom data pipelines utilizing Databricks (PySpark, Spark SQL, and Scala) on AWS infrastructure.

  • Implement and manage a multi-layered Lakehouse architecture (Bronze, Silver, and Gold zones) to curate unstructured contract text, semi-structured logs, and highly structured transactional tables.
  • Architect robust end-to-end data ingestion frameworks supporting high-throughput batch and near real-time data flows from on-premises systems and third-party legal platforms.
  1. Cloud Infrastructure & Platform Optimization
  • Utilize the broad suite of AWS services (including but not limited to S3, Lambda, Glue, EMR, Athena, EC2, and CloudWatch) to support and optimize distributed storage and compute infrastructure.
  • Conduct advanced performance tuning on large-scale Apache Spark workloads optimizing partitioning, indexing, caching strategies, and Databricks cluster utilization to manage cloud run costs efficiently.
  • Automate deployment configurations, orchestrate multi-dependency workflows (via Databricks Jobs/Workflows, Airflow, or Autosys), and build containerized solutions using Docker.
  1. Data Governance, Security & Compliance
  • Enforce strict, fine-grained access controls, row/column-level security, and data classification strategies using Databricks Unity Catalog integrated with AWS IAM and enterprise identity providers.
  • Ensure all data pipelines and lakehouse layers remain strictly compliant with global data privacy regulations (e.g., GDPR) and rigid internal financial audit standards.
  • Implement end-to-end data lineage tracking, validation frameworks, and automated reconciliation routines to preserve absolute data integrity for legal and regulatory reporting.
  1. Downstream Integration & Innovation
  • Collaborate with business analysts and legal operations to expose curated datasets via secure APIs and optimized connectors.
  • Enable seamless consumption of financial and legal analytics through integration with visualization tools like Power BI or automation platforms (Power Apps / Power Automate).
  • Support data readiness for advanced AI/ML models, contract intelligence tools, and eDiscovery search workflows.

Requirements

  • Databricks & Spark: 3+ years of deep, hands-on experience building, scheduling, and debugging data pipelines on Databricks utilizing PySpark, Scala, or Spark SQL.
  • AWS Cloud Suite: Extensive knowledge of AWS core services, with deep familiarity across object storage (S3), serverless compute (Lambda), data cataloging/ETL (Glue), access management (IAM), and encryption (KMS).
  • Data Modeling: Strong proficiency in relational database design, data warehousing structures, schema evolution, and performance tuning techniques (e.g., Delta Lake formats, Apache Iceberg).
  • Programming & Scripting: Strong coding skills in Python and advanced SQL are mandatory.
  • CI/CD & Devops: Proven familiarity with version control (Git) and standard automated deployment workflows., * Regulated Industries: Experience in Financial Services, Asset Management, or handling highly sensitive, audit-driven data environments is highly preferred.
  • Legal Data Concepts: Familiarity with legal data constructs such as contract clauses, corporate matter management, or metadata extraction is a significant advantage.
  • Ownership Mindset: Excellent communication skills, with a track record of collaborating across global, distributed engineering and business architecture teams.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · WWC Europe 2026

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

Videos

See all

Related articles

See all