> Markdown version of [/jobs/ext/1954477-data-engineer](https://www.wearedevelopers.com/jobs/ext/1954477-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** Mizuho Financial Group, Inc. - **Location:** Woodbridge Township, NJ, United States (Remote available) - **Experience:** Starter - **Salary:** $88,000.0 - $105,000.0 - **Contract:** Internship / Graduate position - **Skills:** Abstraction Layers, Artificial Intelligence, Amazon Web Services, Microsoft Azure, Big Data, Cloud Computing, Code Review, Information Systems, Computer Programming, Databases, Data Architecture, Data Deduplication, Information Engineering, Extract Transform Load (ETL), Data Systems, Data Visualization, Relational Databases, Database Queries, Electronic Data Interchange (EDI), Python (Programming Language), Object-Oriented Software Development, Raw Data, Standard Sql, Software Engineering, SQL Databases, SQL Server Integration Services, Cloud Platform System, Data Ingestion, Azure Data Factory, Apache Spark, Git, Data Lakes, Pyspark, Core Data, Information Technology, Apache Kafka, Data Management, Software Version Control, Data Pipelines, Databricks - **Published:** August 6, 2026 - **Apply:** https://mizuho.wd1.myworkdayjobs.com/mizuhoamericas/job/MetroPark/Data-Engineer_R7257-1 ## About the Role * 0-2 years of experience in data engineering, analytics, or a related technical role (internships and academic projects count). * Foundational SQL skills (joins, aggregations, filtering). * Python proficiency with basic OOPS knowledge . * Understanding of core data concepts (tables, schemas, relational data). * Willingness to learn Databricks, Spark, and cloud technologies. * Familiarity with Git or a demonstrated ability to learn version control quickly. Preferred / Nice-to-Have * Exposure to Databricks, Apache Spark, or PySpark (coursework or hands-on). * Awareness of the medallion architecture and Delta Lake basics. * Experience with any cloud platform (Azure, AWS, or GCP). * Relevant coursework, bootcamp, or a Databricks certification (e.g., Data Engineer Associate). * Any experience with data visualization or BI tools., * Eagerness to learn and take feedback. * Attention to detail and care for data accuracy. * Basic problem-solving and logical thinking. * Communicate issues across teams, * Bachelor's degree in Computer Science, Engineering, Information Systems, or related field. Master's degree preferred. * Proven experience in data engineering, software development, or related roles. * Proficiency in programming languages commonly used in data engineering (e.g., Python, Scala, etc.). * Strong knowledge of database systems, data modeling techniques, and SQL proficiency. * Proficiency with ETL tools commonly used in data engineering (e.g., SSIS, Databricks, Azure Data Factory). * Experience with big data technologies and frameworks (e.g., Spark, Kafka, etc.). * Familiarity with cloud platforms and services (e.g., Azure). * Excellent problem-solving skills and attention to detail. * Effective communication and collaboration skills in a team-oriented environment. * Ability to adapt to evolving technologies and business requirements ## Description The IT Data team is responsible for design and development of Data products for the entire firm. The team is embarking on an ambitious new project/implementation. "DEAL". It is the abbreviation for "Data Exchange and Abstraction Layer". It's the next generation, Data Mesh based platform implemented in the Mizuho Azure Cloud on Databricks. Data Mesh is a decentralized data architecture where data is owned and managed by the domain-specific teams that produce the data i.e. Banking, Finance etc., and curate it for downstream consumption. It emphasizes domain-oriented ownership, treating data as a product, providing a self-serve data platform, and using limited federated computational governance from the Data Architecture and Data Management Office. In this role you will be responsible for development of data solutions for the enterprise using innovative and cutting- edge technologies like AI tools / models. The solutions and software developed will be used for reporting and analytics by the entire firm globally. The data solutions developed will have to be accurate, timely and highly scalable. In this role you will be managing large data-sets with complex interdependencies. This is a hands-on software development role. You will be collaborating with teams firmwide to develop solutions. You'll support the development and maintenance of data pipelines on the Databricks Lakehouse platform using the medallion architecture (Bronze/Silver/Gold). You will work under the guidance of senior engineers to ingest, transform, and validate data, growing your skills across the modern data stack., * Assist in building and maintaining ingestion pipelines that land raw data into the Bronze layer. * Support Silver layer transformations under guidance: cleansing, deduplication, and schema enforcement. * Write SQL and PySpark for defined transformation tasks. * Run and monitor scheduled jobs; help investigate and resolve pipeline failures. * Document pipeline logic, transformations, and fixes. * Participate in code reviews as a reviewer-in-training and incorporate feedback on your own work. * Learn team standards for version control, testing, and deployment. ## Related Videos - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Data Governance in the Era of AI](https://www.wearedevelopers.com/videos/1622-data-governance-in-the-era-of-ai) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) ## Related Articles - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)