Lead Data Engineer

Vega Intellisoft Inc.
Foster City, United States of America
2 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Senior

Job location

Foster City, United States of America

Tech stack

API
Agile Methodologies
Artificial Intelligence
Amazon Web Services (AWS)
Amazon Web Services (AWS)
Business Analytics Applications
Confluence
JIRA
Azure
Business Software
Cloud Database
Information Systems
Continuous Integration
Data Architecture
Information Engineering
Data Governance
Data Infrastructure
Data Integration
ETL
Data Warehousing
DevOps
Middleware
Identity and Access Management
Python
Messaging Application Programming Interface
Meta-Data Management
Performance Tuning
Standard Sql
SQL Databases
Data Streaming
Systems Integration
Talend
Unstructured Data
Enterprise Data Management
Data Processing
Azure
Snowflake
Spark
Boomi
Software Troubleshooting
Event Driven Architecture
Microsoft Fabric
Data Lake
PySpark
Information Technology
Kafka
Data Management
Api Design
Cloudwatch
Data Pipelines
Serverless Computing
Mulesoft
Redshift
Databricks

Job description

Hiring Data Engineering Lead with strong expertise in Databricks on AWS, MDM, and enterprise data integration. Lead the design and delivery of modern data platforms that enable trusted, governed, and scalable data consumption across business functions. Experience in cloud-based data engineering, middleware integrations, data governance, and enterprise-scale analytics solutions. Work closely with business, architecture, analytics, and engineering teams to drive data modernization initiatives. This role will be instrumental in enabling the enterprise data modernization journey. Establish a scalable and governed data foundation that supports advanced analytics, AI/ML initiatives, and business decision-making. Success in this role will directly improve data quality, consistency, and accessibility across critical business domains. The architecture and integration patterns defined by this role will serve as a foundation for future data and digital transformation programs., * Lead the design and implementation of Databricks-based data platforms on AWS; Architect scalable Lakehouse solutions supporting enterprise analytics workloads

  • Design and develop complex ETL/ELT pipelines using Databricks, Spark, and Cloud-native services; Drive MDM strategy, implementation, and integration across business applications and data platforms
  • Define data integration patterns using APIs, middleware, event-driven architectures, and messaging frameworks; Establish data governance, metadata management, and data quality frameworks
  • Collaborate with business stakeholders to understand data requirements and translate them into technical solutions; Optimize data processing performance, scalability, and operational monitoring
  • Define CI/CD processes and deployment standards for data engineering assets; Mentor engineering teams and provide technical leadership throughout the project lifecycle
  • Support architecture reviews, solution design discussions, and technical decision-making; Ensure compliance with organizational standards, security requirements, and best practices
  • Facilitate architecture reviews, workshops, and stakeholder discussions; Ability to communicate complex technical concepts to business and executive stakeholders;
  • Stakeholder Management - Build strong relationships with business, IT, and external partners; Manage competing priorities and drive consensus among stakeholders; Demonstrate customer-centric and consultative engagement skills.
  • Leadership Skills - Lead cross-functional and geographically distributed teams; Mentor and guide engineers and junior architects; Influence technical decisions through collaboration
  • Problem Solving & Analytical Thinking - Identify root causes of complex data and integration challenges; Evaluate multiple solution options and recommend optimal approaches; Strong troubleshooting and performance optimization capabilities

Requirements

  • 10+ years of experience in Data Engineering, Data Integration, or Data Platform delivery; hands-on experience with Databricks on AWS; Apache Spark, PySpark, Delta Lake, and Lakehouse architecture
  • Experience designing and implementing enterprise-scale data pipelines; Strong understanding of AWS services such as S3, Glue, Lambda, Redshift, IAM, and CloudWatch
  • Hands-on experience with MDM implementations and integrations; Experience with data quality, data governance, lineage, and master data management processes
  • Strong experience integrating enterprise systems using middleware platforms such as MuleSoft, Boomi, Kafka, or API-based integrations
  • Experience working with structured, semi-structured, and unstructured datasets; Strong SQL and Python development skills
  • Experience with Agile methodologies and DevOps practices; Experience leading distributed teams and managing stakeholder communications
  • Bachelor's Degree or higher in Information Systems, Computer Science, or equivalent experience
  • Skills / Tech Stack Snapshot - Databricks on AWS, Apache Spark, PySpark, Delta Lake, AWS S3, Glue, Lambda, Redshift, MDM Platforms (Informatica MDM, Reltio, Profisee or equivalent), Data Integration & ETL/ELT Frameworks, Middleware Technologies (MuleSoft, Boomi, Kafka, API-led Integrations), Data Warehousing & Data Lake Architecture, Data Governance, MDM, Data Quality & Metadata Management, SQL, Python, Azure DevOps, Jira, Confluence, CI/CD, Agile Delivery, * Experience with Snowflake or Microsoft Fabric; Exposure to AI/ML enablement using Databricks ML or AWS SageMaker
  • Experience with Unity Catalog and data governance frameworks; Data Mesh and Data Product concepts.
  • Experience with real-time streaming using Kafka or Kinesis; Informatica IDMC, Talend, or Azure Data Factory.
  • Knowledge of healthcare, life sciences, retail, manufacturing, or financial services domains.
  • Exposure to GenAI and enterprise AI adoption initiatives

Apply for this position