Junior Data Engineer

Sequoia Connect
United States
25 days ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Internship / Graduate position
Employment type
Full-time (> 32 hours)
Experience level
Starter
Experience required
0 years minimum
Working hours
Regular working hours
Languages
English, Spanish
Job source

Tech stack

Application Programming Interfaces (APIs) Amazon Web Services Amazon S3 Data Analysis Cloud Computing Cloud Database Cloud Engineering Program Optimization Information Engineering Extract Transform Load (ETL) Data Mining JSON
+19 more
Python (Programming Language) Oracle (Applications) Systems Development Life Cycle DataOps Amazon Simple Notification Service (SNS) Software Engineering SQL Databases Data Logging Data Processing Freeform SQL GitHub Copilot Gitlab Pyspark AWS Glue Cloudwatch Amazon Simple Queue Service (SQS) Terraform Software Version Control Data Pipelines

Job description

  • Assist in building and maintaining ETL pipelines using Python and PySpark.
  • Support the development of workflows utilizing AWS Glue, Lambda, and Step Functions.
  • Work extensively with cloud data storage platforms, including S3, Redshift, RDS, and Oracle.
  • Write complex SQL queries for data extraction, transformation, validation, and reporting.
  • Help implement basic monitoring, logging, and error handling for data pipelines.
  • Support the ingestion and processing of data from APIs and JSON payloads.
  • Collaborate with software engineers, data analysts, and business stakeholders to understand requirements.
  • Contribute to code management, technical documentation, and deployment support activities.

Requirements

  • Degree holders for the visa application process.
  • 0 to 4 years of software development or data engineering experience across relevant cloud platforms.
  • Good knowledge of Python and SQL.
  • Solid understanding of AWS services, specifically S3, SNS/SQS, EMR, Glue, Lambda, Redshift, and Step Functions.
  • Strong grasp of ETL concepts and data processing fundamentals.
  • Familiarity with GitLab, Terraform, and the Software Development Life Cycle (SDLC) from development to production.
  • Familiarity with PySpark, AWS managed services, data engineering best practices, and code optimization.
  • High-Performance Mindset: Resilience, emotional intelligence, and a focus on agile delivery.
  • Technologist DNA: A deep understanding of the difference between “coding” and “engineering.”

Desired

  • Exposure to PySpark, Athena, CloudWatch, SNS, and SQS.
  • Internship, project, or academic experience specifically in cloud computing, analytics, or data engineering.
  • Familiarity with cloud-native foundations or AI coding assistants (e.g., GitHub Copilot).

Languages

  • Advanced Oral English: For seamless collaboration with global teams.
  • Advanced Spanish., * 0-4 years of software development experience across the appropriate platform.
  • Good knowledge on Python and SQL.
  • Good understanding to AWS services such as S3, SNS/SQS, EMR, Glue, Lambda, Redshift, and Step Functions.
  • Good understanding of ETL concepts and data processing fundamentals.
  • Familiarity with GitLab/Terraform and SDLC from development to production.
  • Familiarity to PySpark, AWS managed services, data engineering best practices, and code optimization.
  • Good analytical, problem-solving, and communication skills.

Keywords: Phyton, SQL, PySpark, AWS Glue, ETL

About the company

At Sequoia Connect, we are a Talent-First Technology Ecosystem that redefines how elite professionals interact with the global digital landscape. We move beyond traditional models to act as a catalyst for the top 1% of global talent, connecting human potential with complex industrial execution. By joining our inner circle, you are not simply taking a position; you are aligning with a strategic partner dedicated to updating your “Human OS” and accelerating your growth through world-class, high-impact projects.

We are currently partnering with a rapidly growing, automation-led powerhouse that serves 31 Fortune 500 companies across the financial, healthcare, and manufacturing sectors. With a global workforce of over 32,000 employees and a presence in 28 countries, our client is a titan of digital transformation. Their “Automate Everything, Cloudify Everything” strategy ensures you will be working at the absolute forefront of AI-driven automation and cloud solutions.

This is your chance to thrive in a “Customer Success, First and Always” environment that prizes continuous learning and radical ownership. You will collaborate within an international network of expertise across 39 delivery centers worldwide, gaining exposure to complex engineering challenges that redefine industrial standards. If you are a driven professional looking for a dynamic, forward-thinking workplace where your growth is the priority, this is where you belong.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:47 min

Exploring JSON, CBOR, and JOSE for data serialization

Aaron Russell · LIVE

6:14 min

Structuring CI/CD pipelines with integrated security and quality checks

Christoph Ruggenthaler · LIVE

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

1:29 min

Utilizing specific tooling for continuous integration and modeling

Christoph Fassbach Christoph Fassbach +1 · World Congress 2024

2:03 min

Distinguishing type definition constructs from data validation routines

Clemens Vasters Clemens Vasters · World Congress 2025

Videos

See all

Related articles

See all