Data Engineer Generative AI

Hexaware Technologies
Atlanta, United States
17 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Airflow Amazon Web Services Data Analysis Application Integration Architecture Microsoft Azure BigQuery Encodings Data Architecture Data Validation Data Cleansing
+42 more
Information Engineering Data Governance Extract Transform Load (ETL) Data Masking Data Security Data Warehousing Relational Databases DevOps Document-Oriented Databases JSON Python (Programming Language) Machine Learning Meta-Data Management Open Source Technology Cloud Services Search Technologies SQL Databases Unstructured Data Workflow Management Systems Google Cloud Data Ingestion Azure Data Factory Large Language Models Snowflake Prompt Engineering Apache Spark Generative AI Indexer Git Data Lakes Pyspark Kubernetes Information Technology Deployment Automation AWS Glue Integration Frameworks Data Management Data Pipelines Text Files Docker Amazon Redshift Databricks

Job description

We are seeking a skilled Data Engineer with Generative AI experience to design, build, and optimize scalable data platforms that support AI and analytics initiatives. This role will focus on developing reliable data pipelines, preparing high-quality datasets for large language model applications, and enabling secure, production-ready GenAI solutions., * Design, develop, and maintain scalable batch and real-time data pipelines.

  • Build and optimize data ingestion, transformation, validation, and orchestration processes.
  • Develop data models and curated datasets for analytics, machine learning, and Generative AI use cases.
  • Support GenAI applications by preparing, chunking, embedding, indexing, and retrieving enterprise data for Retrieval-Augmented Generation (RAG) workflows.
  • Integrate data sources such as relational databases, APIs, data lakes, document repositories, and streaming platforms.
  • Work with vector databases and embedding models to enable semantic search and GenAI knowledge retrieval.
  • Implement data quality checks, metadata management, lineage, monitoring, and alerting.
  • Partner with data scientists, AI engineers, architects, and business stakeholders to translate requirements into scalable data solutions.
  • Ensure data security, privacy, governance, and access controls are applied across pipelines and AI datasets.
  • Optimize pipeline performance, storage costs, and query efficiency.
  • Document data architecture, pipeline designs, data mappings, and operational procedures.

Requirements

  • Bachelor s degree in Computer Science, Engineering, Data Science, or a related field, or equivalent practical experience.
  • Strong experience in data engineering, ETL/ELT development, and data warehousing.
  • Proficiency in Python and SQL.
  • Experience with data processing frameworks such as Apache Spark, PySpark, Databricks, or similar tools.
  • Experience with cloud data platforms such as AWS, Azure, or Google Cloud Platform.
  • Hands-on experience with data lakes, lakehouses, or cloud warehouses such as Snowflake, Databricks, BigQuery, Redshift, or Synapse.
  • Familiarity with workflow orchestration tools such as Apache Airflow, Azure Data Factory, AWS Glue, or similar platforms.
  • Understanding of Generative AI concepts, including LLMs, embeddings, prompt engineering, vector search, and RAG architectures.
  • Experience integrating APIs and working with semi-structured and unstructured data, including JSON, PDFs, documents, and text files.
  • Strong problem-solving, communication, and collaboration skills., * Experience with vector databases such as Pinecone, Weaviate, Chroma, FAISS, Azure AI Search, or OpenSearch.
  • Experience with GenAI frameworks such as LangChain, LlamaIndex, Semantic Kernel, or similar tools.
  • Familiarity with LLM platforms and services such as Azure OpenAI, Amazon Bedrock, Google Vertex AI, or open-source models.
  • Experience with data governance, cataloging, master data management, and data quality tools.
  • Knowledge of DevOps and CI/CD practices, including Git, Docker, Kubernetes, and automated deployment pipelines.
  • Experience in implementing data masking, PII detection, access controls, and responsible AI practices.
  • Domain experience in financial services, healthcare, retail, or another regulated industry., * Python, SQL, PySpark, Spark
  • ETL/ELT, Data Modeling, Data Warehousing
  • Data Lakes, Lakehouse Architecture, Data Quality
  • Apache Airflow, Databricks, Snowflake
  • AWS, Azure, or Google Cloud Platform
  • LLMs, RAG, Embeddings, Vector Databases
  • API Integration, Semantic Search, Unstructured Data Processing
  • Data Governance, Security, and Privacy

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

3:47 min

Exploring JSON, CBOR, and JOSE for data serialization

Aaron Russell · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

Videos

See all

Related articles

See all