Agentic AI Data Engineer

EXL SERVICE
United States
3 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Compensation
$150,000.0
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services Amazon S3 Continuous Integration Information Engineering Extract Transform Load (ETL) Distributed Computing Environment Python (Programming Language) Machine Learning Standard Sql Azure Machine Learning
+22 more
SQL Databases Unstructured Data AI Infrastructure Data Logging Large Language Models Multi-Agent Systems Prompt Engineering Apache Spark Generative AI Containerization Data Lakes Pyspark Kubernetes Information Technology Apache Kafka Machine Learning Operations Video Streaming Virtual Agents Data Pipelines Docker Amazon Redshift Microservices

Job description

  • Design and implement agentic AI systems that autonomously orchestrate data workflows and decision pipelines
  • Build scalable data pipelines for structured and unstructured data (batch + real-time)
  • Develop and manage LLM-powered applications using retrieval-augmented generation (RAG), tool use, and multi-agent frameworks
  • Integrate AWS AI/ML services into production-grade architectures
  • Develop and optimize data lakes, warehouses, and lakehouse architectures
  • Build APIs and microservices to expose AI/ML capabilities
  • Ensure data quality, governance, and security across pipelines
  • Collaborate with data scientists, ML engineers, and product teams to deploy AI solutions
  • Implement monitoring, logging, and observability for AI agents and pipelinesOptimize cost and performance of cloud-based AI workloads

Requirements

Do you have experience in SQL databases?, Cloud & AWS Ecosystem

  • Strong experience with AWS services, including:
  • Amazon S3, Glue, Lambda, Step Functions
  • Amazon Redshift / Athena
  • Amazon SageMaker (training, deployment, pipelines)
  • Amazon Bedrock (foundation models, agents, knowledge bases)

AI/ML & Agentic Systems

  • Experience with LLMs and generative AI systems
  • Hands-on with agent frameworks (e.g., multi-agent orchestration, tool calling, planning systems)
  • Familiarity with AgentCore / agent orchestration platforms
  • Understanding of RAG architectures , embeddings, and vector databases
  • Experience with model deployment, inference optimization, and prompt engineering

Data Engineering

  • Strong proficiency in Python and SQL
  • Experience with ETL/ELT tools and frameworks
  • Distributed data processing (Spark, PySpark, or similar)
  • Streaming technologies (Kafka, Kinesis, or similar)
  • Data modeling and schema design

Data & AI Infrastructure

  • Experience with vector databases (e.g., Pinecone, FAISS, OpenSearch)
  • Knowledge of data lakehouse architectures (Delta Lake, Iceberg, Hudi)
  • Containerization (Docker) and orchestration (Kubernetes)
  • CI/CD for ML and data pipelines, Cloud & AWS Ecosystem
  • Strong experience with AWS services, including:
  • Amazon S3, Glue, Lambda, Step Functions
  • Amazon Redshift / Athena
  • Amazon SageMaker (training, deployment, pipelines)
  • Amazon Bedrock (foundation models, agents, knowledge bases)

AI/ML & Agentic Systems

  • Experience with LLMs and generative AI systems
  • Hands-on with agent frameworks (e.g., multi-agent orchestration, tool calling, planning systems)
  • Familiarity with AgentCore / agent orchestration platforms
  • Understanding of RAG architectures , embeddings, and vector databases
  • Experience with model deployment, inference optimization, and prompt engineering

Data Engineering

  • Strong proficiency in Python and SQL
  • Experience with ETL/ELT tools and frameworks
  • Distributed data processing (Spark, PySpark, or similar)
  • Streaming technologies (Kafka, Kinesis, or similar)
  • Data modeling and schema design

Data & AI Infrastructure

  • Experience with vector databases (e.g., Pinecone, FAISS, OpenSearch)
  • Knowledge of data lakehouse architectures (Delta Lake, Iceberg, Hudi)
  • Containerization (Docker) and orchestration (Kubernetes)
  • CI/CD for ML and data pipelines

  • Qualifications: Bachelor’s or Master’s degree in Computer Science, Engineering, or related field
  • 4+ years of experience in data engineering or ML engineeringHands-on experience with production-grade AI/ML systems

Benefits & conditions

3.73.7 out of 5 stars United States Hybrid work Up to $150,000 a year - Full-time, This position may pay a base salary of up to $150k per year based on skills and experience.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:43 min

The enduring legacy of the amazon S3 storage API

Chris Heilmann +3 · LIVE

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

5:08 min

Automating data collection and managing crowdsourced training image sets

Kris Howard · LIVE

Videos

See all

Related articles

See all