> Markdown version of [/jobs/ext/804991-aws-databricks-data-engineer](https://www.wearedevelopers.com/jobs/ext/804991-aws-databricks-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AWS Databricks Data Engineer - **Company:** Xoriant Corporation - **Location:** United States (Remote available) - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Airflow, Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, Apache HTTP Server, CA Workload Automation Ae, Cloud Computing, Cloud Database, Program Optimization, Continuous Integration, Data Architecture, Data Governance, Data Integrity, Extract Transform Load (ETL), Data Visualization, Data Warehousing, Relational Databases, Software Debugging, DevOps, Distributed Data Store, Apache Hive, Identity and Access Management, Python (Programming Language), Machine Learning, Automation of Marketing, Meta-Data Management, Performance Tuning, Power BI, Scala (Programming Language), Scripting, Data Classification, Data Ingestion, Microsoft Power Automate, Sql Optimization, Apache Spark, Caching, Indexer, Git, Data Lakes, Pyspark, Data Lineage, Deployment Automation, Real Time Data, Cloudwatch, Software Coding, Software Version Control, Data Pipelines, Serverless Computing, Powerapps, Docker, Databricks - **Published:** June 4, 2026 - **Apply:** https://www.dice.com/job-detail/505dad1b-e988-4291-bc75-80f0a9610dcc ## About the Role * Databricks & Spark: 3+ years of deep, hands-on experience building, scheduling, and debugging data pipelines on Databricks utilizing PySpark, Scala, or Spark SQL. * AWS Cloud Suite: Extensive knowledge of AWS core services, with deep familiarity across object storage (S3), serverless compute (Lambda), data cataloging/ETL (Glue), access management (IAM), and encryption (KMS). * Data Modeling: Strong proficiency in relational database design, data warehousing structures, schema evolution, and performance tuning techniques (e.g., Delta Lake formats, Apache Iceberg). * Programming & Scripting: Strong coding skills in Python and advanced SQL are mandatory. * CI/CD & Devops: Proven familiarity with version control (Git) and standard automated deployment workflows., * Regulated Industries: Experience in Financial Services, Asset Management, or handling highly sensitive, audit-driven data environments is highly preferred. * Legal Data Concepts: Familiarity with legal data constructs such as contract clauses, corporate matter management, or metadata extraction is a significant advantage. * Ownership Mindset: Excellent communication skills, with a track record of collaborating across global, distributed engineering and business architecture teams. ## Description We are seeking a highly skilled Cloud Data Engineer to design, build, and optimize a modern, scalable Legal Data Lakehouse platform. Operating within State Street's Global Technology Services, you will leverage a deep knowledge of the full suite of AWS cloud services combined with high-performance Databricks capabilities to ingest, model, and secure complex enterprise data structures (including contracts, litigation matters, eDiscovery datasets, and global regulatory feeds). This role is critical to establishing a single, highly governed, audit-ready source of truth that powers critical legal operations, compliance analytics, and emerging generative AI/ML use cases across our global footprint., * Design, build, and maintain enterprise-grade, custom data pipelines utilizing Databricks (PySpark, Spark SQL, and Scala) on AWS infrastructure. * Implement and manage a multi-layered Lakehouse architecture (Bronze, Silver, and Gold zones) to curate unstructured contract text, semi-structured logs, and highly structured transactional tables. * Architect robust end-to-end data ingestion frameworks supporting high-throughput batch and near real-time data flows from on-premises systems and third-party legal platforms. 2. Cloud Infrastructure & Platform Optimization * Utilize the broad suite of AWS services (including but not limited to S3, Lambda, Glue, EMR, Athena, EC2, and CloudWatch) to support and optimize distributed storage and compute infrastructure. * Conduct advanced performance tuning on large-scale Apache Spark workloads optimizing partitioning, indexing, caching strategies, and Databricks cluster utilization to manage cloud run costs efficiently. * Automate deployment configurations, orchestrate multi-dependency workflows (via Databricks Jobs/Workflows, Airflow, or Autosys), and build containerized solutions using Docker. 3. Data Governance, Security & Compliance * Enforce strict, fine-grained access controls, row/column-level security, and data classification strategies using Databricks Unity Catalog integrated with AWS IAM and enterprise identity providers. * Ensure all data pipelines and lakehouse layers remain strictly compliant with global data privacy regulations (e.g., GDPR) and rigid internal financial audit standards. * Implement end-to-end data lineage tracking, validation frameworks, and automated reconciliation routines to preserve absolute data integrity for legal and regulatory reporting. 4. Downstream Integration & Innovation * Collaborate with business analysts and legal operations to expose curated datasets via secure APIs and optimized connectors. * Enable seamless consumption of financial and legal analytics through integration with visualization tools like Power BI or automation platforms (Power Apps / Power Automate). * Support data readiness for advanced AI/ML models, contract intelligence tools, and eDiscovery search workflows. ## Related Videos - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) ## Related Articles - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)