AWS Data Engineer

SAVOSH TREE SERVICE INC
Frisco, TX, United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Agile Methodology Amazon Web Services Amazon S3 Data Analysis Bash Shell Big Data Cloud Computing Cloud Database Computer Programming Databases Extract Transform Load (ETL)
+28 more
Data Systems Data Warehousing Database Design Dimensional Modeling Distributed Computing Environment Apache Hadoop Hadoop Distributed File System Apache Hive Python (Programming Language) Unix Shell Microsoft SQL Server Oracle (Applications) Query Optimization Azure Data Lake Shell Script SQL Databases Talend Scripting Cloud Platform System Data Ingestion Apache Spark Data Strategy Pyspark Vba Programming Language AWS Data Analytics Restful APIs Looker Analytics Data Pipelines

Job description

We are seeking a dynamic and highly skilled AWS Data Engineer to join our innovative data team. In this role, you will be at the forefront of designing, developing, and maintaining scalable data pipelines and architectures leveraging Amazon Web Services (AWS) cloud platform. Your expertise will enable the organization to harness the power of big data, facilitate advanced analytics, and support strategic decision-making. This position offers an exciting opportunity to work with cutting-edge technologies, collaborate across teams, and contribute to transformative data solutions that drive business growth., * Design, develop, and optimize scalable ETL (Extract, Transform, Load) pipelines using AWS services such as Glue, S3, Redshift, and Lambda to facilitate efficient data ingestion and processing.

  • Build and maintain robust data models and schemas aligned with best practices in dimensional modeling for data warehousing environments.
  • Develop and implement data management strategies integrating diverse sources including SQL databases (Microsoft SQL Server, Oracle), cloud databases (Azure Data Lake), and big data platforms like Hadoop, Spark, and Apache Hive.
  • Collaborate with cross-functional teams to understand data requirements and translate them into technical solutions using Python, Java, Bash scripting, Shell scripting, and Talend.
  • Ensure high-quality data through rigorous validation, query management, and adherence to security standards within cloud environments.
  • Support model training and analysis activities by providing clean, well-structured datasets; utilize Looker for visualization and reporting purposes.
  • Participate in Agile development processes to continuously improve data workflows and ensure timely delivery of projects.

Requirements

  • Proven experience designing and implementing large-scale data pipelines in cloud environments using AWS services such as S3, Glue, Redshift, Lambda, and related tools.
  • Strong proficiency in SQL programming across multiple database platforms including Microsoft SQL Server, Oracle, and cloud-based systems like Azure Data Lake or similar.
  • Hands-on experience with big data technologies such as Hadoop ecosystem components (HDFS, Hive), Spark (PySpark or Scala), and Apache Hive for distributed data processing.
  • Demonstrated ability to develop ETL workflows using Informatica or Talend; familiarity with RESTful API integration is a plus.
  • Expertise in database design principles including dimensional modeling for data warehousing solutions; experience with query optimization techniques is preferred.
  • Knowledge of scripting languages such as Python, Bash (Unix shell), VBA for automation tasks; experience with analysis skills for deriving insights from complex datasets.
  • Familiarity with Agile methodologies for project management; ability to adapt quickly in a fast-paced environment while maintaining attention to detail.

Join us if you’re passionate about leveraging cloud technology to transform raw data into actionable insights! Bring your expertise in AWS cloud platforms combined with your skills in big data processing, database design, and analytics to make a meaningful impact on our organization’s success.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

2:10 min

Why organizations combine big data and machine learning

Ayon Roy · LIVE

1:43 min

AWS infrastructure stack and data flow pipeline overview

Artem Volk Artem Volk +1 · WWC 2024

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

Videos

See all

Related articles

See all