Data Engineer

Tiger Analytics
Jersey City, NJ, United States
8 days ago
Apply on www.juju.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Airflow Amazon Web Services Amazon S3 Data Analysis Automated Storage and Retrieval Systems Data Architecture Data Cleansing Data Infrastructure Data Integration Extract Transform Load (ETL) Data Systems
+18 more
Machine Learning SQL Databases Unstructured Data Data Processing System Availability Large Language Models Apache Spark Generative AI AWS Lambda Git Data Lakes Infrastructure Automation Frameworks AWS Glue Machine Learning Operations Software Version Control Data Pipelines Amazon Redshift Databricks

Job description

We are seeking an experienced Data Engineer to join our data team. In this role, you will be responsible for designing, building, and maintaining scalable data pipelines, data integration processes, and data infrastructure on AWS cloud. You will collaborate closely with data scientists, analysts, and AI teams to support analytics, machine learning, and Generative AI initiatives across the organization., * Design, develop, and deploy end-to-end data pipelines on AWS cloud infrastructure using services such as Amazon S3, AWS Glue, AWS Lambda, Amazon Redshift, etc.

  • Implement data processing and transformation workflows using Databricks, Apache Spark, and SQL to support analytics, reporting, and AI-driven use cases.
  • Build and maintain orchestration workflows using Apache Airflow to automate data pipeline execution, scheduling, and monitoring.
  • Support data preparation and ingestion for AI/ML and Generative AI workloads, including handling structured and unstructured datasets.
  • Enable data pipelines that support LLM-based applications, vector embeddings, and knowledge retrieval systems.
  • Lead the migration of legacy data systems to modern cloud-based data architectures.
  • Develop and maintain CI/CD pipelines for data workflows and platform automation.
  • Collaborate with data scientists, ML engineers, and AI teams to ensure data availability for model training, inference, and GenAI applications.
  • Optimize data pipelines for performance, reliability, scalability, and cost-effectiveness using AWS best practices.

Requirements

  • 8+ years experience as a Data Engineer working with AWS cloud services.
  • Hands-on experience with AWS services such as S3, Glue, Lambda, Redshift, and related data platform tools.
  • Experience building data pipelines using Databricks, Apache Spark, and SQL.
  • Experience with Apache Airflow for workflow orchestration.
  • Strong understanding of data modeling, data lake/lakehouse architectures, and ETL/ELT frameworks.
  • Experience with CI/CD pipelines and version control systems (Git).
  • Exposure to Generative AI or LLM-based applications.
  • Experience supporting data pipelines for AI/ML workloads.
  • Familiarity with vector databases, embeddings, and Retrieval-Augmented Generation (RAG) architectures.
  • Experience working with LLM APIs or AI frameworks such as LangChain.
  • Understanding of MLOps workflows and model deployment pipelines.

Benefits & conditions

Significant career development opportunities exist as the company grows. The position offers a unique opportunity to be part of a small, challenging, and entrepreneurial environment, with a high degree of individual responsibility.

About the company

Tiger Analytics is a fast-growing advanced analytics consulting firm. Our consultants bring deep expertise in Data Science, Machine Learning and AI. We are the trusted analytics partner for several Fortune 100 companies, enabling them to generate business value from data. Our business value and leadership has been recognized by various market research firms, including Forrester and Gartner. We are looking for top-notch talent as we continue to build the best analytics global consulting team in the world.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.juju.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

55 sec

Validating data processing architectures via containerized events

Modood Alvi · World Congress 2025

2:19 min

Introduction to Apache Airflow for advanced orchestration

Alan Mazankiewicz · LIVE

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

1:11 min

Deploying and running Airflow in cloud environments

Alan Mazankiewicz · LIVE

1:59 min

Evolving roles in AI driven software teams

Ignacio Riesgo Ignacio Riesgo +1 · World Congress 2024

Videos

See all

Related articles

See all