> Markdown version of [/jobs/ext/2972980-data-engineer-generative-ai](https://www.wearedevelopers.com/jobs/ext/2972980-data-engineer-generative-ai). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer Generative AI - **Company:** Hexaware Technologies - **Location:** Atlanta, United States - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Airflow, Amazon Web Services, Data Analysis, Application Integration Architecture, Microsoft Azure, BigQuery, Encodings, Data Architecture, Data Validation, Data Cleansing, Information Engineering, Data Governance, Extract Transform Load (ETL), Data Masking, Data Security, Data Warehousing, Relational Databases, DevOps, Document-Oriented Databases, JSON, Python (Programming Language), Machine Learning, Meta-Data Management, Open Source Technology, Cloud Services, Search Technologies, SQL Databases, Unstructured Data, Workflow Management Systems, Google Cloud, Data Ingestion, Azure Data Factory, Large Language Models, Snowflake, Prompt Engineering, Apache Spark, Generative AI, Indexer, Git, Data Lakes, Pyspark, Kubernetes, Information Technology, Deployment Automation, AWS Glue, Integration Frameworks, Data Management, Data Pipelines, Text Files, Docker, Amazon Redshift, Databricks - **Published:** September 18, 2026 - **Apply:** https://www.dice.com/job-detail/bdbdfc7a-5a87-463f-9fe5-2da153a4f4e1 ## About the Role * Bachelor s degree in Computer Science, Engineering, Data Science, or a related field, or equivalent practical experience. * Strong experience in data engineering, ETL/ELT development, and data warehousing. * Proficiency in Python and SQL. * Experience with data processing frameworks such as Apache Spark, PySpark, Databricks, or similar tools. * Experience with cloud data platforms such as AWS, Azure, or Google Cloud Platform. * Hands-on experience with data lakes, lakehouses, or cloud warehouses such as Snowflake, Databricks, BigQuery, Redshift, or Synapse. * Familiarity with workflow orchestration tools such as Apache Airflow, Azure Data Factory, AWS Glue, or similar platforms. * Understanding of Generative AI concepts, including LLMs, embeddings, prompt engineering, vector search, and RAG architectures. * Experience integrating APIs and working with semi-structured and unstructured data, including JSON, PDFs, documents, and text files. * Strong problem-solving, communication, and collaboration skills., * Experience with vector databases such as Pinecone, Weaviate, Chroma, FAISS, Azure AI Search, or OpenSearch. * Experience with GenAI frameworks such as LangChain, LlamaIndex, Semantic Kernel, or similar tools. * Familiarity with LLM platforms and services such as Azure OpenAI, Amazon Bedrock, Google Vertex AI, or open-source models. * Experience with data governance, cataloging, master data management, and data quality tools. * Knowledge of DevOps and CI/CD practices, including Git, Docker, Kubernetes, and automated deployment pipelines. * Experience in implementing data masking, PII detection, access controls, and responsible AI practices. * Domain experience in financial services, healthcare, retail, or another regulated industry., * Python, SQL, PySpark, Spark * ETL/ELT, Data Modeling, Data Warehousing * Data Lakes, Lakehouse Architecture, Data Quality * Apache Airflow, Databricks, Snowflake * AWS, Azure, or Google Cloud Platform * LLMs, RAG, Embeddings, Vector Databases * API Integration, Semantic Search, Unstructured Data Processing * Data Governance, Security, and Privacy ## Description We are seeking a skilled Data Engineer with Generative AI experience to design, build, and optimize scalable data platforms that support AI and analytics initiatives. This role will focus on developing reliable data pipelines, preparing high-quality datasets for large language model applications, and enabling secure, production-ready GenAI solutions., * Design, develop, and maintain scalable batch and real-time data pipelines. * Build and optimize data ingestion, transformation, validation, and orchestration processes. * Develop data models and curated datasets for analytics, machine learning, and Generative AI use cases. * Support GenAI applications by preparing, chunking, embedding, indexing, and retrieving enterprise data for Retrieval-Augmented Generation (RAG) workflows. * Integrate data sources such as relational databases, APIs, data lakes, document repositories, and streaming platforms. * Work with vector databases and embedding models to enable semantic search and GenAI knowledge retrieval. * Implement data quality checks, metadata management, lineage, monitoring, and alerting. * Partner with data scientists, AI engineers, architects, and business stakeholders to translate requirements into scalable data solutions. * Ensure data security, privacy, governance, and access controls are applied across pipelines and AI datasets. * Optimize pipeline performance, storage costs, and query efficiency. * Document data architecture, pipeline designs, data mappings, and operational procedures. ## Related Videos - [Beyond GPT: Building Unified GenAI Platforms for the Enterprise of Tomorrow](https://www.wearedevelopers.com/videos/1525-beyond-gpt-building-unified-genai-platforms-for-the-enterprise-of-tomorrow) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Tips and Tricks for Working with JSON](https://www.wearedevelopers.com/videos/1229-tips-and-tricks-for-working-with-json) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) ## Related Articles - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)