AI Data Engineer

Propertyvalue Quantum Technologies Llc
Tallahassee, United States of America
yesterday

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Senior
Compensation
$ 187K

Job location

Tallahassee, United States of America

Tech stack

Artificial Intelligence
Airflow
Amazon Web Services (AWS)
Apache HTTP Server
Azure
Cloud Computing
Data Architecture
Data Validation
Information Engineering
Data Governance
ETL
Data Mining
Relational Databases
Database Testing
Hive
Python
Machine Learning
Metadata
Query Optimization
Search Technologies
SQL Databases
Data Streaming
Unstructured Data
Workflow Management Systems
Google Cloud Platform
Retrieval-Augmented Generation
Large Language Models
Spark
Generative AI
Event Driven Architecture
Data Lake
PySpark
Kafka
Machine Learning Operations
Data Delivery
GPT
Data Pipelines
Docker
Databricks

Job description

  • Design, develop, and optimize scalable batch and real-time data pipelines using Apache Spark (PySpark/Spark SQL).
  • Write complex, high-performance SQL queries for data extraction, transformation, analytics, and reporting.
  • Build and maintain ETL/ELT pipelines to ingest, cleanse, transform, and integrate structured and unstructured data.
  • Prepare, curate, and validate datasets for Machine Learning and Generative AI applications.
  • Develop and optimize RAG (Retrieval-Augmented Generation) data pipelines using vector databases and document processing frameworks.
  • Integrate enterprise data with Large Language Models (LLMs) such as OpenAI GPT, Azure OpenAI, Claude, or Gemini.
  • Implement AI-powered data quality validation, anomaly detection, and automated monitoring solutions.
  • Perform data validation, testing, and quality assurance to ensure data accuracy, completeness, and consistency.
  • Optimize Spark jobs, SQL queries, and distributed processing for maximum performance and scalability.
  • Collaborate with Data Scientists, AI Engineers, Business Analysts, and Application Developers to support AI initiatives.
  • Monitor ETL workflows and AI data pipelines to ensure reliable and timely data delivery.
  • Maintain technical documentation for data architecture, AI pipelines, metadata, and data models.
  • Follow best practices for data governance, security, compliance, and AI model lifecycle management.

Requirements

  • 5+ years of experience as a Data Engineer.
  • Strong hands-on experience with SQL and query optimization.
  • Extensive experience with Apache Spark (PySpark/Spark SQL).
  • Strong Python programming experience.
  • Experience building and maintaining scalable ETL/ELT pipelines.
  • Strong understanding of relational databases and data modeling.
  • Experience with Generative AI, LLMs, and AI-powered data engineering workflows.
  • Knowledge of RAG architecture, embeddings, vector databases (Pinecone, ChromaDB, FAISS, or Weaviate), and semantic search.
  • Experience with AI frameworks such as LangChain or LlamaIndex.
  • Experience with cloud platforms such as AWS, Azure, or Google Cloud Platform.
  • Experience with data testing, validation, and quality assurance.
  • Strong analytical, troubleshooting, and communication skills.
  • Ability to work effectively in an onsite, collaborative environment.
  • Experience with Databricks.
  • Experience with Delta Lake, Apache Iceberg, or Apache Hudi.
  • Knowledge of Kafka, event-driven architectures, or streaming data pipelines.
  • Experience with Airflow or other workflow orchestration tools.
  • Experience with Docker and Kubernetes.
  • Familiarity with MLOps tools such as MLflow.
  • Experience implementing enterprise AI governance and responsible AI practices.

Apply for this position