> Markdown version of [/jobs/ext/917412-data-analytics-engineering-data-engineer-iii](https://www.wearedevelopers.com/jobs/ext/917412-data-analytics-engineering-data-engineer-iii). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Analytics & Engineering - Data Engineer III - **Company:** Generative Ai - **Location:** Menlo Park, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Training Data, Artificial Intelligence, Airflow, Data Cleansing, Information Engineering, Extract Transform Load (ETL), Cursor (Graphical User Interface Elements), Database Queries, Software Debugging, Fault Tolerance, Machine Learning, Object Detection, Operational Databases, Query Optimization, SQL Databases, Data Streaming, Sql Optimization, Large Language Models, Prompt Engineering, Generative AI, Indexer, Information Technology, Machine Learning Operations, Data Pipelines - **Published:** June 2, 2026 - **Apply:** https://www.dice.com/job-detail/c76da7f5-4a0f-4255-b042-62d95c4bd3e3 ## About the Role Advanced SQL & data pipeline expertise. Complex queries, query optimization, pipeline orchestration frameworks (Airflow, Dataswarm, or equivalent). Experience integrating ML models into data pipelines. Calling inference endpoints, managing model versions, batching requests, handling inference failures at scale. Proficiency with AI-assisted coding agents (e.g., Copilot, Cursor, Codex). Expected to leverage AI tools as a force multiplier for writing, debugging, and reviewing code, building pipelines faster, and accelerating day-to-day engineering workflows Strong verbal and written communication skills, problem-solving ability, and cross-functional collaboration. Preferred Working knowledge of embeddings and vector representations like generating, storing, indexing, and querying embeddings (FAISS, Milvus, or equivalent). Familiarity with content-understanding models like image classifiers, object detection, OCR, NSFW detection, aesthetic scoring. Experience with LLMs for data tasks like prompt engineering for annotation, data cleaning, or evaluation using LLM APIs. Knowledge of generative AI like diffusion models, image generation, evaluation metrics (FID, CLIP score, etc.)., Bachelor''s degree or higher in Computer Science, Data Engineering, Machine Learning, or a related STEM field. 5+ years of industry experience in data engineering, ML engineering, or a hybrid role involving both data pipelines and model serving/inference. Demonstrated track record of building and operating production data pipelines that invoke ML models at scale. Previous experience at Meta is preferred but not required. ## Description This role sits at the intersection of Data Engineering and ML Systems. The Senior AI Data Engineer will own end-to-end data pipelines that don''t just move and transform data, but enrich it through remote model inference, managing the systems complexity of async execution, capacity allocation, retry/fallback logic, and throughput optimization that comes with it. This is not a pure ETL-with-SQL role; it demands hands-on systems experience with distributed inference infrastructure. Our team develops comprehensive data curation and evaluation solutions for image generation models across quality dimensions including visual quality, prompt adherence, identity preservation, naturalness, and visual text generation., AI-Augmented Data Pipelines: Design and maintain AI-augmented, large-scale data pipelines (billions of images) integrating traditional transformations with ML models (classifiers, embeddings, LLMs) for cleaning and annotation. Remote Inference Orchestration: Own the systems for remote ML model inference orchestration within pipelines, managing batching, retries, async jobs, and ensuring graceful degradation. Feature Pipelines: Build and maintain scalable pipelines for generating, storing, and serving vector embeddings, including nearest-neighbor index management and quality validation. Data Curation at Scale: Source, filter, and curate training datasets using a combination of SQL and model-derived signals (e.g., aesthetic scores, NSFW classifiers), owning the end-to-end data flow and maintaining governance, quality, and compliance. Additional Responsibilities LLM-Assisted Annotation: Design and operate pipelines that use LLMs and vision models for automated annotation of training data, including auditing workflows to measure and improve annotation model performance. Tooling & Frameworks: Contribute to shared tooling and frameworks that make it easier for the broader team to build AI-augmented data pipelines - e.g., reusable operators for model invocation, standard patterns for async job management. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Optimizing Discovery: PostgreSQL's Role in Transforming GetYourGuide's Search](https://www.wearedevelopers.com/videos/1647-optimizing-discovery-postgresql-s-role-in-transforming-getyourguide-s-search) - [Fireside Chat: Deep Learning, Deep Impact: Harnessing AI for Language Innovation](https://www.wearedevelopers.com/videos/612-fireside-chat-deep-learning-deep-impact-harnessing-ai-for-language-innovation) - [Developer Experience, Platform Engineering and AI powered Apps](https://www.wearedevelopers.com/videos/990-developer-experience-platform-engineering-and-ai-powered-apps) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) - [Exploring 5 Key Applications of AI Abundance with Blockchain Assurance](https://www.wearedevelopers.com/videos/971-exploring-5-key-applications-of-ai-abundance-with-blockchain-assurance) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it)