Junior Data Engineer
Role details
Job location
Tech stack
Job description
- Design, develop, and optimize large-scale ETL (Extract, Transform, Load) pipelines using tools like Informatica, Talend, or custom scripting with Python and Bash to ensure reliable data flow across systems.
- Architect and maintain cloud-based data solutions utilizing AWS, Azure Data Lake, and other public cloud platforms to support scalable data warehousing and analytics.
- Collaborate with cross-functional teams to understand business needs and translate them into effective data models using dimensional modeling techniques for data warehouses.
- Develop and implement advanced SQL queries for query management, reporting, and business intelligence tools such as Looker.
- Manage big data systems including Hadoop, Apache Hive, Spark, and related technologies to process vast datasets efficiently.
- Ensure data quality, security, and compliance by establishing best practices for data management integration and governance.
- Support model training and analysis efforts by providing clean, well-structured datasets for machine learning initiatives.
Requirements
We are seeking a dynamic and highly skilled Junior Data Engineer to join our innovative data team. In this role, you will lead the design, development, and management of complex data systems that empower data-driven decision-making across the organization. You will leverage your expertise in big data technologies, cloud platforms, and data modeling to build scalable, efficient, and secure data pipelines and warehouses. Your contributions will be instrumental in transforming raw data into actionable insights, supporting strategic initiatives, and driving business growth., * Proven experience as a Data Engineer or similar role with a strong background in software development and data engineering principles.
- Extensive knowledge of cloud databases such as Azure Data Lake, AWS services (e.g., S3), and other public cloud environments.
- Hands-on experience with big data systems including Hadoop ecosystem components like Hive and Spark for large-scale data processing.
- Proficiency in SQL programming along with Python scripting for automation and pipeline development; familiarity with VBA or Shell Scripting is a plus.
- Strong understanding of data modeling concepts including dimensional modeling for data warehousing design.
- Experience working with business intelligence tools such as Looker or similar platforms to facilitate analytics reporting.
- Familiarity with ETL pipeline development using tools like Talend or Informatica; knowledge of RESTful APIs for integrating external systems is advantageous.
- Ability to work within Agile teams while managing multiple priorities effectively; excellent analysis skills to interpret complex datasets.
Join us to be at the forefront of transforming raw data into strategic assets! We are committed to fostering an energetic environment where your expertise drives innovation and growth across our organization.
Benefits & conditions
Pulled from the full job description 401(k) 401(k) matching Paid time off Vision insurance Dental insurance Life insurance, * 401(k)
- 401(k) matching
- Dental insurance
- Life insurance
- Paid time off
- Vision insurance