Data Engineer Generative AI
Hexaware Technologies
Atlanta, United States
17 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.dice.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source
Tech stack
Application Programming Interfaces (APIs)
Artificial Intelligence
Airflow
Amazon Web Services
Data Analysis
Application Integration Architecture
Microsoft Azure
BigQuery
Encodings
Data Architecture
Data Validation
Data Cleansing
+42 more
Information Engineering
Data Governance
Extract Transform Load (ETL)
Data Masking
Data Security
Data Warehousing
Relational Databases
DevOps
Document-Oriented Databases
JSON
Python (Programming Language)
Machine Learning
Meta-Data Management
Open Source Technology
Cloud Services
Search Technologies
SQL Databases
Unstructured Data
Workflow Management Systems
Google Cloud
Data Ingestion
Azure Data Factory
Large Language Models
Snowflake
Prompt Engineering
Apache Spark
Generative AI
Indexer
Git
Data Lakes
Pyspark
Kubernetes
Information Technology
Deployment Automation
AWS Glue
Integration Frameworks
Data Management
Data Pipelines
Text Files
Docker
Amazon Redshift
Databricks
Job description
We are seeking a skilled Data Engineer with Generative AI experience to design, build, and optimize scalable data platforms that support AI and analytics initiatives. This role will focus on developing reliable data pipelines, preparing high-quality datasets for large language model applications, and enabling secure, production-ready GenAI solutions., * Design, develop, and maintain scalable batch and real-time data pipelines.
- Build and optimize data ingestion, transformation, validation, and orchestration processes.
- Develop data models and curated datasets for analytics, machine learning, and Generative AI use cases.
- Support GenAI applications by preparing, chunking, embedding, indexing, and retrieving enterprise data for Retrieval-Augmented Generation (RAG) workflows.
- Integrate data sources such as relational databases, APIs, data lakes, document repositories, and streaming platforms.
- Work with vector databases and embedding models to enable semantic search and GenAI knowledge retrieval.
- Implement data quality checks, metadata management, lineage, monitoring, and alerting.
- Partner with data scientists, AI engineers, architects, and business stakeholders to translate requirements into scalable data solutions.
- Ensure data security, privacy, governance, and access controls are applied across pipelines and AI datasets.
- Optimize pipeline performance, storage costs, and query efficiency.
- Document data architecture, pipeline designs, data mappings, and operational procedures.
Requirements
- Bachelor s degree in Computer Science, Engineering, Data Science, or a related field, or equivalent practical experience.
- Strong experience in data engineering, ETL/ELT development, and data warehousing.
- Proficiency in Python and SQL.
- Experience with data processing frameworks such as Apache Spark, PySpark, Databricks, or similar tools.
- Experience with cloud data platforms such as AWS, Azure, or Google Cloud Platform.
- Hands-on experience with data lakes, lakehouses, or cloud warehouses such as Snowflake, Databricks, BigQuery, Redshift, or Synapse.
- Familiarity with workflow orchestration tools such as Apache Airflow, Azure Data Factory, AWS Glue, or similar platforms.
- Understanding of Generative AI concepts, including LLMs, embeddings, prompt engineering, vector search, and RAG architectures.
- Experience integrating APIs and working with semi-structured and unstructured data, including JSON, PDFs, documents, and text files.
- Strong problem-solving, communication, and collaboration skills., * Experience with vector databases such as Pinecone, Weaviate, Chroma, FAISS, Azure AI Search, or OpenSearch.
- Experience with GenAI frameworks such as LangChain, LlamaIndex, Semantic Kernel, or similar tools.
- Familiarity with LLM platforms and services such as Azure OpenAI, Amazon Bedrock, Google Vertex AI, or open-source models.
- Experience with data governance, cataloging, master data management, and data quality tools.
- Knowledge of DevOps and CI/CD practices, including Git, Docker, Kubernetes, and automated deployment pipelines.
- Experience in implementing data masking, PII detection, access controls, and responsible AI practices.
- Domain experience in financial services, healthcare, retail, or another regulated industry., * Python, SQL, PySpark, Spark
- ETL/ELT, Data Modeling, Data Warehousing
- Data Lakes, Lakehouse Architecture, Data Quality
- Apache Airflow, Databricks, Snowflake
- AWS, Azure, or Google Cloud Platform
- LLMs, RAG, Embeddings, Vector Databases
- API Integration, Semantic Search, Unstructured Data Processing
- Data Governance, Security, and Privacy
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.dice.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
ER
Erin Rifkin
over 1 year ago
MH
Michael Hunger
Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?
9 months ago
LM
Luis Minvielle
How to Become an AI Engineer
almost 3 years ago
BB
Benedikt Bischof
Making Data Warehouses Fast: A Developer’s Story
about 4 years ago
CH
Chris Heilmann
Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production
about 2 years ago
BB
Benedikt Bischof
MLOps And AI Driven Development
over 4 years ago