Data Engineer (AWS)- IND

Insight Global
St. Louis Park, MN, United States
about 2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Cloud Computing Encodings Continuous Integration Information Engineering Extract Transform Load (ETL) Python (Programming Language) Search Technologies Software Construction Unstructured Data
+8 more
AI Infrastructure Data Logging Enterprise Software Applications Large Language Models Generative AI AWS Lambda Data Pipelines Serverless Computing

Job description

Insight Global is seeking an AI Data Engineer with deep expertise in AWS to join one of our medical device clients local to India. In this role, you will develop scalable ETL/ELT pipelines to ingest and process large volumes of unstructured data, leveraging cloud-native services such as Azure Document Intelligence, AWS Textract, and other OCR tools to extract and prepare content at scale. You will architect and manage high-quality embedding pipelines, chunking strategies, and vector database integrations-including platforms like Azure AI Search-to support Retrieval-Augmented Generation (RAG) workflows and intelligent search capabilities. You will build retrieval and orchestration pipelines that connect data to LLMs, implement resilient CI/CD workflows, and ensure strong logging, monitoring, and error handling for production reliability. Working closely with platform and application teams, you will integrate LLM-powered features into enterprise applications and deploy services using functions, containers, and event-driven cloud technologies such as AWS Lambda and Azure Functions. A strong focus on security, compliance, scalability, and performance will be critical as you help advance the organization’s AI engineering ecosystem.

Requirements

5+ years in data engineering or AI infrastructure roles

-Expertise in AWS

-Hands-on experience with vector stores and embedding pipelines

-Strong Python development experience

-Experience with OCR/document intelligence tools

-Strong familiarity with LLMs, RAG architectures, embeddings, and retrieval techniques

-Experience with CICD and software engineering best practices

-Experience building and maintaining data pipelines

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on juju.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:17 min

Optimizing character encoding with Kim variable byte encoding

Douglas Crockford Douglas Crockford · WWC 2024

2:45 min

Executing event-driven code using AWS Lambda functions

Sebastien Stormacq Sebastien Stormacq · LIVE

1:10 min

Exposing sensitive information through partial search logs

Dennis Schulz Dennis Schulz +1 · WWC Europe 2026

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

4:12 min

Distilling cross-encoder models into smaller efficient sentence embedding models

Marek Suppa · LIVE

Videos

See all

Related articles

See all