Sr. Databricks Engineer - Hybrid NYC

Randstad
New York, NY, United States
12 days ago
Apply on www.randstadusa.com
Prepare application

Role details

Contract type
Contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Compensation
$145,600.0 - $156,000.0
Working hours
Regular working hours

Tech stack

Java (Programming Language) Artificial Intelligence Airflow Amazon Web Services Amazon Elastic Compute Cloud Amazon S3 Applications Architecture Big Data Bioinformatics C++ (Programming Language) Cloud Computing Cloud Database
+60 more
Cloud Engineering Information Systems Computer Engineering Continuous Integration Data Architecture Information Engineering Data Governance Data Infrastructure Dataspaces Data Systems Cursor (Graphical User Interface Elements) Software Design Patterns Distributed Computing Environment Distributed Systems Fault Tolerance Apache Hadoop Apache HBase Apache Hive Identity and Access Management Data Intelligence IntelliJ IDEA Python (Programming Language) Machine Learning Meta-Data Management Apache Oozie Productivity Software Query Optimization Standard Sql DataOps Shell Script Amazon Simple Notification Service (SNS) Software Engineering SQL Databases Sqoop Data Streaming Tableau (Software) Technical Data Management Systems Data Storage Technologies Data Ingestion Large Language Models Snowflake Apache Spark Electronic Medical Records Generative AI Git Event Driven Architecture Data Lakes Pyspark Information Technology Low Latency Apache Kafka Spark Streaming Data Management Virtual Agents Amazon Simple Queue Service (SQS) Splunk Data Pipelines Docker Databricks Programming Languages

Job description

job summary: We are seeking an experienced Senior Data Engineer to design, develop, and optimize scalable data platforms that power enterprise analytics, machine learning, and AI-driven solutions. The ideal candidate will bring deep expertise in modern data engineering practices, cloud-native architectures, Lakehouse platforms, and distributed data processing technologies.

This role will play a critical part in building reliable, high-performance data ecosystems leveraging Databricks, Spark, Delta Lake, Snowflake, Kafka, and AWS, while also contributing to the adoption of Generative AI, Large Language Models (LLMs), Retrieval Augmented Generation (RAG), and AI-assisted engineering solutions.

Key Responsibilities

Data Platform Engineering

Design, build, and maintain scalable batch and real-time data pipelines.

Develop and optimize data ingestion, transformation, and processing frameworks for structured, semi-structured, and unstructured datasets.

Implement modern Lakehouse architectures utilizing Databricks, Delta Lake, and Medallion (Bronze, Silver, Gold) design patterns.

Build data solutions that support enterprise analytics, reporting, and machine learning initiatives.

Ensure data quality, governance, lineage, security, and compliance across data ecosystems.

Big Data & Streaming Solutions

Develop distributed data processing applications using PySpark and Spark SQL.

Build and maintain streaming pipelines using Kafka, Spark Structured Streaming, and AWS Kinesis.

Design fault-tolerant, scalable systems capable of processing large data volumes with low latency.

Optimize workload performance through partitioning strategies, clustering, caching, and query tuning.

Cloud & Lakehouse Architecture

Architect and implement cloud-based data solutions on AWS.

Utilize AWS services including S3, EMR, EC2, Athena, Redshift, RDS, Lambda, IAM, SNS, and SQS.

Design data storage and processing strategies that maximize reliability while minimizing operational costs.

Support migration initiatives from traditional Hadoop and EMR environments to modern cloud-native platforms.

Data Operations & Automation

Develop orchestration and scheduling frameworks using Airflow and Databricks Workflows.

Build CI/CD pipelines and automation frameworks for deployment, monitoring, and data platform operations.

Collaborate closely with architects, analysts, data scientists, and business stakeholders to deliver enterprise-grade solutions.

AI & Intelligent Platform Engineering

Implement Generative AI-powered solutions for engineering productivity and operational excellence.

Develop applications leveraging Large Language Models (LLMs), Retrieval Augmented Generation (RAG), Vector Databases, and Model Context Protocol (MCP).

Build AI-assisted documentation, developer productivity tooling, and intelligent platform capabilities.

Evaluate emerging AI technologies and identify opportunities for adoption within data engineering processes.

Required Qualifications

Bachelor’s or Master’s degree in Computer Engineering, Computer Science, Information Systems, or a related field.

8+ years of experience in software engineering, data engineering, or big data platform development.

Strong experience designing and implementing enterprise-scale data pipelines.

Hands-on expertise with:

Python

PySpark

Spark SQL

SQL

Kafka

Databricks

Delta Lake

Snowflake

Hive

Experience building data solutions on AWS cloud platforms.

Strong understanding of distributed computing, data modeling, and large-scale data processing.

Experience with Git-based development workflows and CI/CD practices.

Excellent analytical, troubleshooting, and problem-solving skills.

Preferred Qualifications

Experience with real-time streaming architectures and event-driven systems.

Knowledge of data governance, metadata management, and data quality frameworks.

Experience with generative AI technologies including:

LLMs

RAG

Vector Databases

AI Agents

MCP integrations

Experience developing developer productivity tools and AI-assisted engineering workflows.

Exposure to enterprise supply chain, retail, healthcare, or manufacturing data domains.

AWS certifications are highly preferred.

Technical Skills

Programming Languages

Python

Java

SQL

Shell Scripting

C/C++

Big Data & Data Engineering

PySpark

Spark SQL

Hive

Databricks

Delta Lake

Snowflake

Kafka

HBase

Sqoop

Workflow & Orchestration

Apache Airflow

Databricks Workflows

Oozie

Cloud Technologies

AWS S3

EMR

EC2

Athena

Redshift

RDS

IAM

Lambda

SNS

SQS

AI & Modern Engineering

Generative AI

Large Language Models (LLMs)

Retrieval Augmented Generation (RAG)

Agentic AI Systems

Model Context Protocol (MCP)

Vector Databases

Visualization & Tools

Tableau

Git

Docker

Splunk

IntelliJ IDEA

PyCharm

Cursor

Preferred Certifications

AWS Certified Solutions Architect - Associate

AWS Certified Cloud Practitioner

Databricks Certifications (preferred)

What Success Looks Like

Deliver highly scalable and reliable data pipelines.

Improve platform performance, efficiency, and cost optimization.

Enable enterprise-wide analytics and AI initiatives through trusted data products.

Drive modernization of data platforms and adoption of cloud-native architectures.

Leverage AI technologies to enhance engineering efficiency, automation, and innovation.

Ideal Candidate Profile: A senior-level data engineer with extensive experience in Databricks, Spark, AWS, Kafka, Snowflake, and Lakehouse architectures, who is equally passionate about modern AI technologies and building intelligent data platforms for the future.

location: New York, New York job type: Contract salary: $70 - 75 per hour work hours: 9am to 6pm education: Bachelors

responsibilities: We are seeking an experienced Senior Data Engineer to design, develop, and optimize scalable data platforms that power enterprise analytics, machine learning, and AI-driven solutions. The ideal candidate will bring deep expertise in modern data engineering practices, cloud-native architectures, Lakehouse platforms, and distributed data processing technologies.

This role will play a critical part in building reliable, high-performance data ecosystems leveraging Databricks, Spark, Delta Lake, Snowflake, Kafka, and AWS , while also contributing to the adoption of Generative AI, Large Language Models (LLMs), Retrieval Augmented Generation (RAG), and AI-assisted engineering solutions .

Key Responsibilities

Data Platform Engineering

  • Design, build, and maintain scalable batch and real-time data pipelines.
  • Develop and optimize data ingestion, transformation, and processing frameworks for structured, semi-structured, and unstructured datasets.
  • Implement modern Lakehouse architectures utilizing Databricks, Delta Lake, and Medallion (Bronze, Silver, Gold) design patterns.
  • Build data solutions that support enterprise analytics, reporting, and machine learning initiatives.
  • Ensure data quality, governance, lineage, security, and compliance across data ecosystems.

Big Data & Streaming Solutions

  • Develop distributed data processing applications using PySpark and Spark SQL.
  • Build and maintain streaming pipelines using Kafka, Spark Structured Streaming, and AWS Kinesis.
  • Design fault-tolerant, scalable systems capable of processing large data volumes with low latency.
  • Optimize workload performance through partitioning strategies, clustering, caching, and query tuning.

Cloud & Lakehouse Architecture

  • Architect and implement cloud-based data solutions on AWS.
  • Utilize AWS services including S3, EMR, EC2, Athena, Redshift, RDS, Lambda, IAM, SNS, and SQS.
  • Design data storage and processing strategies that maximize reliability while minimizing operational costs.
  • Support migration initiatives from traditional Hadoop and EMR environments to modern cloud-native platforms.

Data Operations & Automation

  • Develop orchestration and scheduling frameworks using Airflow and Databricks Workflows.
  • Build CI/CD pipelines and automation frameworks for deployment, monitoring, and data platform operations.
  • Collaborate closely with architects, analysts, data scientists, and business stakeholders to deliver enterprise-grade solutions.

AI & Intelligent Platform Engineering

  • Implement Generative AI-powered solutions for engineering productivity and operational excellence.
  • Develop applications leveraging Large Language Models (LLMs), Retrieval Augmented Generation (RAG), Vector Databases, and Model Context Protocol (MCP).
  • Build AI-assisted documentation, developer productivity tooling, and intelligent platform capabilities.
  • Evaluate emerging AI technologies and identify opportunities for adoption within data engineering processes.

Required Qualifications

  • Bachelor’s or Master’s degree in Computer Engineering, Computer Science, Information Systems, or a related field.
  • 8+ years of experience in software engineering, data engineering, or big data platform development.
  • Strong experience designing and implementing enterprise-scale data pipelines.
  • Hands-on expertise with:
  • Python
  • PySpark
  • Spark SQL
  • SQL
  • Kafka
  • Databricks
  • Delta Lake
  • Snowflake
  • Hive
  • Experience building data solutions on AWS cloud platforms.
  • Strong understanding of distributed computing, data modeling, and large-scale data processing.
  • Experience with Git-based development workflows and CI/CD practices.
  • Excellent analytical, troubleshooting, and problem-solving skills.

Preferred Qualifications

  • Experience with real-time streaming architectures and event-driven systems.
  • Knowledge of data governance, metadata management, and data quality frameworks.
  • Experience with generative AI technologies including:
  • LLMs
  • RAG
  • Vector Databases
  • AI Agents
  • MCP integrations
  • Experience developing developer productivity tools and AI-assisted engineering workflows.
  • Exposure to enterprise supply chain, retail, healthcare, or manufacturing data domains.
  • AWS certifications are highly preferred.

Technical Skills

Programming Languages

  • Python
  • Java
  • SQL
  • Shell Scripting
  • C/C++

Big Data & Data Engineering

  • PySpark
  • Spark SQL
  • Hive
  • Databricks
  • Delta Lake
  • Snowflake
  • Kafka
  • HBase
  • Sqoop

Workflow & Orchestration

  • Apache Airflow
  • Databricks Workflows
  • Oozie

Cloud Technologies

  • AWS S3
  • EMR
  • EC2
  • Athena
  • Redshift
  • RDS
  • IAM
  • Lambda
  • SNS
  • SQS

AI & Modern Engineering

  • Generative AI
  • Large Language Models (LLMs)
  • Retrieval Augmented Generation (RAG)
  • Agentic AI Systems
  • Model Context Protocol (MCP)
  • Vector Databases

Visualization & Tools

  • Tableau
  • Git
  • Docker
  • Splunk
  • IntelliJ IDEA
  • PyCharm
  • Cursor

Preferred Certifications

  • AWS Certified Solutions Architect - Associate
  • AWS Certified Cloud Practitioner
  • Databricks Certifications (preferred)

What Success Looks Like

  • Deliver highly scalable and reliable data pipelines.
  • Improve platform performance, efficiency, and cost optimization.
  • Enable enterprise-wide analytics and AI initiatives through trusted data products.
  • Drive modernization of data platforms and adoption of cloud-native architectures.
  • Leverage AI technologies to enhance engineering efficiency, automation, and innovation.

Ideal Candidate Profile: A senior-level data engineer with extensive experience in Databricks, Spark, AWS, Kafka, Snowflake, and Lakehouse architectures, who is equally passionate about modern AI technologies and building intelligent data platforms for the future.

qualifications: We are seeking an experienced Senior Data Engineer to design, develop, and optimize scalable data platforms that power enterprise analytics, machine learning, and AI-driven solutions. The ideal candidate will bring deep expertise in modern data engineering practices, cloud-native architectures, Lakehouse platforms, and distributed data processing technologies.

This role will play a critical part in building reliable, high-performance data ecosystems leveraging Databricks, Spark, Delta Lake, Snowflake, Kafka, and AWS, while also contributing to the adoption of Generative AI, Large Language Models (LLMs), Retrieval Augmented Generation (RAG), and AI-assisted engineering solutions.

Key Responsibilities

Data Platform Engineering

Design, build, and maintain scalable batch and real-time data pipelines.

Develop and optimize data ingestion, transformation, and processing frameworks for structured, semi-structured, and unstructured datasets.

Implement modern Lakehouse architectures utilizing Databricks, Delta Lake, and Medallion (Bronze, Silver, Gold) design patterns.

Build data solutions that support enterprise analytics, reporting, and machine learning initiatives.

Ensure data quality, governance, lineage, security, and compliance across data ecosystems.

Big Data & Streaming Solutions

Develop distributed data processing applications using PySpark and Spark SQL.

Build and maintain streaming pipelines using Kafka, Spark Structured Streaming, and AWS Kinesis.

Design fault-tolerant, scalable systems capable of processing large data volumes with low latency.

Optimize workload performance through partitioning strategies, clustering, caching, and query tuning.

Cloud & Lakehouse Architecture

Architect and implement cloud-based data solutions on AWS.

Utilize AWS services including S3, EMR, EC2, Athena, Redshift, RDS, Lambda, IAM, SNS, and SQS.

Design data storage and processing strategies that maximize reliability while minimizing operational costs.

Support migration initiatives from traditional Hadoop and EMR environments to modern cloud-native platforms.

Data Operations & Automation

Develop orchestration and scheduling frameworks using Airflow and Databricks Workflows.

Build CI/CD pipelines and automation frameworks for deployment, monitoring, and data platform operations.

Collaborate closely with architects, analysts, data scientists, and business stakeholders to deliver enterprise-grade solutions.

AI & Intelligent Platform Engineering

Implement Generative AI-powered solutions for engineering productivity and operational excellence.

Develop applications leveraging Large Language Models (LLMs), Retrieval Augmented Generation (RAG), Vector Databases, and Model Context Protocol (MCP).

Build AI-assisted documentation, developer productivity tooling, and intelligent platform capabilities.

Evaluate emerging AI technologies and identify opportunities for adoption within data engineering processes.

Required Qualifications

Bachelor’s or Master’s degree in Computer Engineering, Computer Science, Information Systems, or a related field.

8+ years of experience in software engineering, data engineering, or big data platform development.

Strong experience designing and implementing enterprise-scale data pipelines.

Hands-on expertise with:

Python

PySpark

Spark SQL

SQL

Kafka

Databricks

Delta Lake

Snowflake

Hive

Experience building data solutions on AWS cloud platforms.

Strong understanding of distributed computing, data modeling, and large-scale data processing.

Experience with Git-based development workflows and CI/CD practices.

Excellent analytical, troubleshooting, and problem-solving skills.

Preferred Qualifications

Experience with real-time streaming architectures and event-driven systems.

Knowledge of data governance, metadata management, and data quality frameworks.

Experience with generative AI technologies including:

LLMs

RAG

Vector Databases

AI Agents

MCP integrations

Experience developing developer productivity tools and AI-assisted engineering workflows.

Exposure to enterprise supply chain, retail, healthcare, or manufacturing data domains.

AWS certifications are highly preferred.

Technical Skills

Programming Languages

Python

Java

SQL

Shell Scripting

C/C++

Big Data & Data Engineering

PySpark

Spark SQL

Hive

Databricks

Delta Lake

Snowflake

Kafka

HBase

Sqoop

Workflow & Orchestration

Apache Airflow

Databricks Workflows

Oozie

Cloud Technologies

AWS S3

EMR

EC2

Athena

Redshift

RDS

IAM

Lambda

SNS

SQS

AI & Modern Engineering

Generative AI

Large Language Models (LLMs)

Retrieval Augmented Generation (RAG)

Agentic AI Systems

Model Context Protocol (MCP)

Vector Databases

Visualization & Tools

Tableau

Git

Docker

Splunk

IntelliJ IDEA

PyCharm

Cursor

Preferred Certifications

AWS Certified Solutions Architect - Associate

AWS Certified Cloud Practitioner

Databricks Certifications (preferred)

What Success Looks Like

Deliver highly scalable and reliable data pipelines.

Improve platform performance, efficiency, and cost optimization.

Enable enterprise-wide analytics and AI initiatives through trusted data products.

Drive modernization of data platforms and adoption of cloud-native architectures.

Leverage AI technologies to enhance engineering efficiency, automation, and innovation.

Ideal Candidate Profile: A senior-level data engineer with extensive experience in Databricks, Spark, AWS, Kafka, Snowflake, and Lakehouse architectures, who is equally passionate about modern AI technologies and building intelligent data platforms for the future.

skills: Airflow,Apache Airflow,EC2,S3,AWS S3,SNS,SQS,AWS,AWS services,AWS cloud,HBase,Hadoop,Spark SQL,Hive,Kafka,Oozie,Spark,Develop applications,AI-driven solutions,AI,AI technologies,Big Data,large data,large-scale data processing,C/C++,Cloud,Cloud Technologies,cloud-based data,cloud-native architectures,Computer Engineering,CI/CD,Cursor,Lakehouse Architecture,data governance,data quality frameworks,Data Platform Engineering,data platform,big data platform,data ingestion,Delta Lake,data platforms,real-time data pipelines,data pipelines,data storage,Streaming,AWS Kinesis,real-time streaming,data solutions,Data Operations,Databricks,data ecosystems,distributed data processing,distributed computing,Docker,EMR,event-driven systems,fault-tolerant,Generative AI,Retrieval Augmented Generation (RAG),generative AI technologies,Git,IAM,data engineering,Information Systems,Computer Science,IntelliJ IDEA,Java,Large Language Models,LLMs,low latency,machine learning,metadata management,productivity tools,Programming Languages,PySpark,Python,query tuning,SQL,Shell Scripting,Snowflake,design patterns,software engineering,Spark Structured Streaming,Splunk,Sqoop,Tableau,manufacturing data,Agentic AI,analytical,troubleshooting,problem-solving skills,reliability,high-performance,AWS certifications,Automation,enterprise analytics,Big Data,cloud-native architectures,Java,SQL,cost optimization,data platforms,data modeling,data quality,data products,governance,healthcare,innovation,operational excellence,Platform Engineering,retail,security,scalable systems,scheduling,supply chain,Vector Databases,Visualization,Workflows

Equal Opportunity Employer: Race, Color, Religion, Sex, Sexual Orientation, Gender Identity, National Origin, Age, Genetic Information, Disability, Protected Veteran Status, or any other legally protected group status.

At Randstad Digital, we welcome people of all abilities and want to ensure that our hiring and interview process meets the needs of all applicants. If you require a reasonable accommodation to make your application or interview experience a great one, please contact HRsupport@randstadusa.com.

Pay offered to a successful candidate will be based on several factors including the candidate’s education, work experience, work location, specific job duties, certifications, etc. In addition, Randstad Digital offers a comprehensive benefits package, including: medical, prescription, dental, vision, AD&D, and life insurance offerings, short-term disability, and a 401K plan (all benefits are based on eligibility).

This posting is open for thirty (30) days.

,

We are seeking an experienced Senior Data Engineer to design, develop, and optimize scalable data platforms that power enterprise analytics, machine learning, and AI-driven solutions. The ideal candidate will bring deep expertise in modern data engineering practices, cloud-native architectures, Lakehouse platforms, and distributed data processing technologies.

This role will play a critical part in building reliable, high-performance data ecosystems leveraging Databricks, Spark, Delta Lake, Snowflake, Kafka, and AWS, while also contributing to the adoption of Generative AI, Large Language Models (LLMs), Retrieval Augmented Generation (RAG), and AI-assisted engineering solutions.

Key Responsibilities

Data Platform Engineering

  • Design, build, and maintain scalable batch and real-time data pipelines.
  • Develop and optimize data ingestion, transformation, and processing frameworks for structured, semi-structured, and unstructured datasets.
  • Implement modern Lakehouse architectures utilizing Databricks, Delta Lake, and Medallion (Bronze, Silver, Gold) design patterns.
  • Build data solutions that support enterprise analytics, reporting, and machine learning initiatives.
  • Ensure data quality, governance, lineage, security, and compliance across data ecosystems.

Big Data & Streaming Solutions

  • Develop distributed data processing applications using PySpark and Spark SQL.
  • Build and maintain streaming pipelines using Kafka, Spark Structured Streaming, and AWS Kinesis.
  • Design fault-tolerant, scalable systems capable of processing large data volumes with low latency.
  • Optimize workload performance through partitioning strategies, clustering, caching, and query tuning.

Cloud & Lakehouse Architecture

  • Architect and implement cloud-based data solutions on AWS.
  • Utilize AWS services including S3, EMR, EC2, Athena, Redshift, RDS, Lambda, IAM, SNS, and SQS.
  • Design data storage and processing strategies that maximize reliability while minimizing operational costs.
  • Support migration initiatives from traditional Hadoop and EMR environments to modern cloud-native platforms.

Data Operations & Automation

  • Develop orchestration and scheduling frameworks using Airflow and Databricks Workflows.
  • Build CI/CD pipelines and automation frameworks for deployment, monitoring, and data platform operations.
  • Collaborate closely with architects, analysts, data scientists, and business stakeholders to deliver enterprise-grade solutions.

AI & Intelligent Platform Engineering

  • Implement Generative AI-powered solutions for engineering productivity and operational excellence.
  • Develop applications leveraging Large Language Models (LLMs), Retrieval Augmented Generation (RAG), Vector Databases, and Model Context Protocol (MCP).
  • Build AI-assisted documentation, developer productivity tooling, and intelligent platform capabilities.
  • Evaluate emerging AI technologies and identify opportunities for adoption within data engineering processes.

Required Qualifications

  • Bachelor’s or Master’s degree in Computer Engineering, Computer Science, Information Systems, or a related field.
  • 8+ years of experience in software engineering, data engineering, or big data platform development.
  • Strong experience designing and implementing enterprise-scale data pipelines.
  • Hands-on expertise with:
  • Python
  • PySpark
  • Spark SQL
  • SQL
  • Kafka
  • Databricks
  • Delta Lake
  • Snowflake
  • Hive
  • Experience building data solutions on AWS cloud platforms.
  • Strong understanding of distributed computing, data modeling, and large-scale data processing.
  • Experience with Git-based development workflows and CI/CD practices.
  • Excellent analytical, troubleshooting, and problem-solving skills.

Preferred Qualifications

  • Experience with real-time streaming architectures and event-driven systems.
  • Knowledge of data governance, metadata management, and data quality frameworks.
  • Experience with generative AI technologies including:
  • LLMs
  • RAG
  • Vector Databases
  • AI Agents
  • MCP integrations
  • Experience developing developer productivity tools and AI-assisted engineering workflows.
  • Exposure to enterprise supply chain, retail, healthcare, or manufacturing data domains.
  • AWS certifications are highly preferred.

Technical Skills

Programming Languages

  • Python
  • Java
  • SQL
  • Shell Scripting
  • C/C++

Big Data & Data Engineering

  • PySpark
  • Spark SQL
  • Hive
  • Databricks
  • Delta Lake
  • Snowflake
  • Kafka
  • HBase
  • Sqoop

Workflow & Orchestration

  • Apache Airflow
  • Databricks Workflows
  • Oozie

Cloud Technologies

  • AWS S3
  • EMR
  • EC2
  • Athena
  • Redshift
  • RDS
  • IAM
  • Lambda
  • SNS
  • SQS

AI & Modern Engineering

  • Generative AI
  • Large Language Models (LLMs)
  • Retrieval Augmented Generation (RAG)
  • Agentic AI Systems
  • Model Context Protocol (MCP)
  • Vector Databases

Visualization & Tools

  • Tableau
  • Git
  • Docker
  • Splunk
  • IntelliJ IDEA
  • PyCharm
  • Cursor

Preferred Certifications

  • AWS Certified Solutions Architect - Associate
  • AWS Certified Cloud Practitioner
  • Databricks Certifications (preferred)

What Success Looks Like

  • Deliver highly scalable and reliable data pipelines.
  • Improve platform performance, efficiency, and cost optimization.
  • Enable enterprise-wide analytics and AI initiatives through trusted data products.
  • Drive modernization of data platforms and adoption of cloud-native architectures.
  • Leverage AI technologies to enhance engineering efficiency, automation, and innovation.

Ideal Candidate Profile: A senior-level data engineer with extensive experience in Databricks, Spark, AWS, Kafka, Snowflake, and Lakehouse architectures, who is equally passionate about modern AI technologies and building intelligent data platforms for the future.

Requirements

Airflow,Apache Airflow,EC2,S3,AWS S3,SNS,SQS,AWS,AWS services,AWS cloud,HBase,Hadoop,Spark SQL,Hive,Kafka,Oozie,Spark,Develop applications,AI-driven solutions,AI,AI technologies,Big Data,large data,large-scale data processing,C/C++,Cloud,Cloud Technologies,cloud-based data,cloud-native architectures,Computer Engineering,CI/CD,Cursor,Lakehouse Architecture,data governance,data quality frameworks,Data Platform Engineering,data platform,big data platform,data ingestion,Delta Lake,data platforms,real-time data pipelines,data pipelines,data storage,Streaming,AWS Kinesis,real-time streaming,data solutions,Data Operations,Databricks,data ecosystems,distributed data processing,distributed computing,Docker,EMR,event-driven systems,fault-tolerant,Generative AI,Retrieval Augmented Generation (RAG),generative AI technologies,Git,IAM,data engineering,Information Systems,Computer Science,IntelliJ IDEA,Java,Large Language Models,LLMs,low latency,machine learning,metadata management,productivity tools,Programming Languages,PySpark,Python,query tuning,SQL,Shell Scripting,Snowflake,design patterns,software engineering,Spark Structured Streaming,Splunk,Sqoop,Tableau,manufacturing data,Agentic AI,analytical,troubleshooting,problem-solving skills,reliability,high-performance,AWS certifications,Automation,enterprise analytics,Big Data,cloud-native architectures,Java,SQL,cost optimization,data platforms,data modeling,data quality,data products,governance,healthcare,innovation,operational excellence,Platform Engineering,retail,security,scalable systems,scheduling,supply chain,Vector Databases,Visualization,Workflows

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.randstadusa.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:41 min

Visualizing the complex developer journey for JVM ecosystems

Bobur Umurzokov · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

2:34 min

Capabilities of the Apache Spark processing engine

Ayon Roy · LIVE

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

Videos

See all

Related articles

See all