> Markdown version of [/jobs/ext/1277789-data-engineer-with-java](https://www.wearedevelopers.com/jobs/ext/1277789-data-engineer-with-java). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer with Java - **Company:** Amazon.com, Inc. - **Location:** Alpharetta, GA, United States - **Experience:** Expert - **Salary:** $101,500.0 - $169,100.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Automation of Tests, Microsoft Azure, Cloud Database, Extract Transform Load (ETL), Data Transformation, Data Migration, Electronic Data Interchange (EDI), Python (Programming Language), PostgreSQL, Machine Learning, NoSQL, Performance Tuning, Query Optimization, Release Management, Power BI, SQL Databases, Data Streaming, Enterprise Software Applications, Data Ingestion, Cloud Monitoring, Snowflake, Grafana, Data Lakes, Pyspark, Gitlab-ci, Deployment Automation, Cassandra, Data Analytics, Enterprise Integration, Apache Kafka, Video Streaming, Stream Analytics, Data Pipelines, Databricks - **Published:** July 15, 2026 - **Apply:** https://www.careerjet.com/jobad/us99badf8601c36d49108096615ca9f050 ## About the Role * Strong experience in Python and PySpark development. * Hands-on experience with Azure Databricks and databricks SQL. * Experience in Java-based streaming and ingestion frameworks. * Strong knowledge of Apache Kafka streaming concepts. * Experience working with PostgreSQL databases. * Experience with YugabyteDB or Cassandra-based NoSQL databases. * Strong SQL development and query optimization skills. * Hands-on experience with GitLab CI/CD pipeline development and deployment automation. * Understanding cloud-based data engineering and distributed processing concepts. * Experience in data migration projects, especially Snowflake on-prem to Azure cloud migration. * Experience designing enterprise-scale data lake or lakehouse architectures. * Knowledge of streaming architectures and real-time analytics. * Familiarity with cloud monitoring and observability tools. Key Skills: Azure, Databricks, Java, Python, Pyspark. ## Description * The ideal candidate should possess strong expertise in Databricks, PySpark, Python, Java-based streaming technologies, GitLab CI/CD pipelines, and cloud migration initiatives., * Design and develop scalable Databricks ETL/ELT pipelines (Lakeflow & LakeBase) using Azure Databricks, PySpark, and Python. * Implement real-time and batch data ingestion frameworks using Kafka and Java-based streaming solutions. * Develop and optimize data processing workflows in Azure Databricks. * Integrate and manage data movement between PostgreSQL, YugabyteDB (Cassandra-based NoSQL), and cloud platforms. * Build reusable frameworks for data ingestion, transformation, validation, and orchestration. * Develop SQL-based data transformations, reporting datasets, and performance optimization solutions. * Design and implement GitLab CI/CD pipelines for automated deployment, testing, and release management of Databricks notebooks, jobs, and data pipelines. * Support Snowflake on-premises to Azure cloud migration initiatives. * Ensure coding standards, performance tuning, monitoring, and operational stability of data pipelines. * Develop Power BI dashboards and reports for business intelligence and analytics reporting. * Develop API automation and integration solutions for data exchange between enterprise systems., + $101,500-169,100 per year Are you passionate about enabling analytics teams to make fast, confident, data-driven decisions? Do you enjoy building and optimizing the foundational data assets that power busin… + 3 days ago, We are seeking a visionary Machine Learning Engineer Lead to spearhead our experimental ML initiatives and drive innovation across the organization. This role combines technical le… + 3 days ago + ## Related Videos - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Leveraging Real time data in FSIs](https://www.wearedevelopers.com/videos/806-leveraging-real-time-data-in-fsis) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) ## Related Articles - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Data Analyst Salary in the UK](https://www.wearedevelopers.com/magazine/278-data-analyst-salary-in-the-uk)