Data Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+28 more
Job description
The IT Data team is responsible for design and development of Data products for the entire firm. The team is embarking on an ambitious new project/implementation. “DEAL”. It is the abbreviation for “Data Exchange and Abstraction Layer”. It’s the next generation, Data Mesh based platform implemented in the Mizuho Azure Cloud on Databricks. Data Mesh is a decentralized data architecture where data is owned and managed by the domain-specific teams that produce the data i.e. Banking, Finance etc., and curate it for downstream consumption. It emphasizes domain-oriented ownership, treating data as a product, providing a self-serve data platform, and using limited federated computational governance from the Data Architecture and Data Management Office.
In this role you will be responsible for development of data solutions for the enterprise using innovative and cutting- edge technologies like AI tools / models. The solutions and software developed will be used for reporting and analytics by the entire firm globally. The data solutions developed will have to be accurate, timely and highly scalable. In this role you will be managing large data-sets with complex interdependencies. This is a hands-on software development role. You will be collaborating with teams firmwide to develop solutions.
You’ll support the development and maintenance of data pipelines on the Databricks Lakehouse platform using the medallion architecture (Bronze/Silver/Gold). You will work under the guidance of senior engineers to ingest, transform, and validate data, growing your skills across the modern data stack., * Assist in building and maintaining ingestion pipelines that land raw data into the Bronze layer.
- Support Silver layer transformations under guidance: cleansing, deduplication, and schema enforcement.
- Write SQL and PySpark for defined transformation tasks.
- Run and monitor scheduled jobs; help investigate and resolve pipeline failures.
- Document pipeline logic, transformations, and fixes.
- Participate in code reviews as a reviewer-in-training and incorporate feedback on your own work.
- Learn team standards for version control, testing, and deployment.
Requirements
- 0-2 years of experience in data engineering, analytics, or a related technical role (internships and academic projects count).
- Foundational SQL skills (joins, aggregations, filtering).
- Python proficiency with basic OOPS knowledge .
- Understanding of core data concepts (tables, schemas, relational data).
- Willingness to learn Databricks, Spark, and cloud technologies.
- Familiarity with Git or a demonstrated ability to learn version control quickly.
Preferred / Nice-to-Have
- Exposure to Databricks, Apache Spark, or PySpark (coursework or hands-on).
- Awareness of the medallion architecture and Delta Lake basics.
- Experience with any cloud platform (Azure, AWS, or GCP).
- Relevant coursework, bootcamp, or a Databricks certification (e.g., Data Engineer Associate).
- Any experience with data visualization or BI tools., * Eagerness to learn and take feedback.
- Attention to detail and care for data accuracy.
- Basic problem-solving and logical thinking.
- Communicate issues across teams, * Bachelor’s degree in Computer Science, Engineering, Information Systems, or related field. Master’s degree preferred.
- Proven experience in data engineering, software development, or related roles.
- Proficiency in programming languages commonly used in data engineering (e.g., Python, Scala, etc.).
- Strong knowledge of database systems, data modeling techniques, and SQL proficiency.
- Proficiency with ETL tools commonly used in data engineering (e.g., SSIS, Databricks, Azure Data Factory).
- Experience with big data technologies and frameworks (e.g., Spark, Kafka, etc.).
- Familiarity with cloud platforms and services (e.g., Azure).
- Excellent problem-solving skills and attention to detail.
- Effective communication and collaboration skills in a team-oriented environment.
- Ability to adapt to evolving technologies and business requirements
Benefits & conditions
The expected base salary ranges from $88,000 - $105,000. Salary offers are based on a wide range of factors including relevant skills, training, experience, education, and, where applicable, certifications and licenses obtained. Market and organizational factors are also considered. In addition to salary and a generous employee benefits package, including Medical, Dental and 401K plans, successful candidates are also eligible to receive a discretionary bonus.
About the company
Mizuho Financial Group, Inc. is the 15th largest bank in the world as measured by total assets of ~$2 trillion. Mizuho’s 60,000 employees worldwide offer comprehensive financial services to clients in 35 countries and 800 offices throughout the Americas, EMEA and Asia. Mizuho Americas is a leading provider of corporate and investment banking services to clients in the US, Canada, and Latin America. Through its acquisition of Greenhill , Mizuho provides M&A, restructuring and private capital advisory capabilities across Americas, Europe and Asia. Mizuho Americas employs approximately 3,500 professionals, and its capabilities span corporate and investment banking, capital markets, equity and fixed income sales & trading, derivatives, FX, custody and research. Visit www.mizuhoamericas.com.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on mizuho.wd1.myworkdayjobs.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Top Big Data Technologies That You Need to Know
Making Data Warehouses Fast: A Developer’s Story
How to Become an AI Engineer
The Most Popular IT Jobs on the Market