Data Engineer with Java

Amazon.com, Inc.
Alpharetta, GA, United States
27 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Compensation
$101,500.0 - $169,100.0
Working hours
Regular working hours

Tech stack

Java (Programming Language) Automation of Tests Microsoft Azure Cloud Database Extract Transform Load (ETL) Data Transformation Data Migration Electronic Data Interchange (EDI) Python (Programming Language) PostgreSQL Machine Learning NoSQL
+23 more
Performance Tuning Query Optimization Release Management Power BI SQL Databases Data Streaming Enterprise Software Applications Data Ingestion Cloud Monitoring Snowflake Grafana Data Lakes Pyspark Gitlab-ci Deployment Automation Cassandra Data Analytics Enterprise Integration Apache Kafka Video Streaming Stream Analytics Data Pipelines Databricks

Job description

  • The ideal candidate should possess strong expertise in Databricks, PySpark, Python, Java-based streaming technologies, GitLab CI/CD pipelines, and cloud migration initiatives., * Design and develop scalable Databricks ETL/ELT pipelines (Lakeflow & LakeBase) using Azure Databricks, PySpark, and Python.
  • Implement real-time and batch data ingestion frameworks using Kafka and Java-based streaming solutions.
  • Develop and optimize data processing workflows in Azure Databricks.
  • Integrate and manage data movement between PostgreSQL, YugabyteDB (Cassandra-based NoSQL), and cloud platforms.
  • Build reusable frameworks for data ingestion, transformation, validation, and orchestration.
  • Develop SQL-based data transformations, reporting datasets, and performance optimization solutions.
  • Design and implement GitLab CI/CD pipelines for automated deployment, testing, and release management of Databricks notebooks, jobs, and data pipelines.
  • Support Snowflake on-premises to Azure cloud migration initiatives.
  • Ensure coding standards, performance tuning, monitoring, and operational stability of data pipelines.
  • Develop Power BI dashboards and reports for business intelligence and analytics reporting.
  • Develop API automation and integration solutions for data exchange between enterprise systems., + $101,500-169,100 per year Are you passionate about enabling analytics teams to make fast, confident, data-driven decisions? Do you enjoy building and optimizing the foundational data assets that power busin…

  • 3 days ago, We are seeking a visionary Machine Learning Engineer Lead to spearhead our experimental ML initiatives and drive innovation across the organization. This role combines technical le…
  • 3 days ago +

Requirements

  • Strong experience in Python and PySpark development.
  • Hands-on experience with Azure Databricks and databricks SQL.
  • Experience in Java-based streaming and ingestion frameworks.
  • Strong knowledge of Apache Kafka streaming concepts.
  • Experience working with PostgreSQL databases.
  • Experience with YugabyteDB or Cassandra-based NoSQL databases.
  • Strong SQL development and query optimization skills.
  • Hands-on experience with GitLab CI/CD pipeline development and deployment automation.
  • Understanding cloud-based data engineering and distributed processing concepts.
  • Experience in data migration projects, especially Snowflake on-prem to Azure cloud migration.
  • Experience designing enterprise-scale data lake or lakehouse architectures.
  • Knowledge of streaming architectures and real-time analytics.
  • Familiarity with cloud monitoring and observability tools.

Key Skills: Azure, Databricks, Java, Python, Pyspark.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:37 min

Comparing traditional SQL tables versus NoSQL non-tabular databases

Stanimira Vlaeva · JS Congress

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:16 min

Terminology differences between relational and NoSQL databases

Tim Faulkes · LIVE

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · WWC Europe 2026

Videos

See all

Related articles

See all