AI Data Engineer

United IT Solutions
United States
4 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Cloud Engineering Continuous Integration Data Validation Extract Transform Load (ETL) DevOps Distributed Computing Environment Distributed Systems Github Machine Learning Open Source Technology
+13 more
Performance Tuning Systems Integration Data Processing Large Language Models Apache Spark Generative AI Containerization Pyspark Kubernetes Data Pipelines Docker Jenkins Databricks

Job description

· Pipeline Engineering: Design, build, and maintain production-level data pipelines to deploy and operationalize ML and LLM workflows.

· LLM & RAG Integration: Implement Retrieval-Augmented Generation (RAG) frameworks using libraries like LangChain or LlamaIndex to query structured and unstructured data sources.

· API & System Integration: Integrate LLM APIs (e.g., OpenAI, Anthropic, or open-source models) into data processing workflows.

· Performance Optimization: Optimize distributed workloads, data processing engines, and pipeline latency for real-time and batch execution.

· CI/CD & DevOps: Build and maintain CI/CD pipelines to deploy data and AI workflows using relevant SDKs and automation tools.

· Data Quality & Validation: Implement strict schema validation rules and data quality checks to ensure reliable pipeline execution.

· AI Evaluation & Quality Control: Monitor and measure output quality using key metrics such as retrieval quality, answer correctness, and faithfulness.

Requirements

· Experience: Proven experience as a Data Engineer building production-grade ETL/ELT data pipelines.

· LLM / AI Concepts: Minimum working knowledge of ML concepts, LLM architectures, vector databases, and RAG frameworks (e.g., LangChain, LlamaIndex).

· API Integration: Hands-on experience integrating third-party or self-hosted LLM APIs into data pipelines.

· Distributed Computing: Experience optimizing distributed data processing workloads (e.g., PySpark, Spark, Databricks, Ray, or Cloud-native processing services).

· CI/CD & Automation: Solid understanding of CI/CD pipeline implementation, deployment SDKs, and containerization (e.g., Docker, Kubernetes, GitHub Actions, Jenkins).

· Schema & Data Quality: Expertise in enforcing schema validation rules, data contracts, and pipeline performance optimization.

· Evaluation Metrics: Familiarity with AI/RAG evaluation metrics (e.g., retrieval precision, answer correctness, context relevance, faithfulness).

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

1:59 min

Evolving roles in AI driven software teams

Ignacio Riesgo Ignacio Riesgo +1 · World Congress 2024

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

Videos

See all

Related articles

See all