> Markdown version of [/jobs/ext/657768-data-engineer-ai-llm-expertise](https://www.wearedevelopers.com/jobs/ext/657768-data-engineer-ai-llm-expertise). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer (AI & LLM Expertise) - **Company:** Pyramid Consulting Inc. - **Location:** Plano, TX, United States - **Salary:** $114,400.0 - $124,800.0 - **Contract:** Temporary contract - **Skills:** Artificial Intelligence, Amazon Web Services, Microsoft Azure, Client Access Licensing, Cloud Computing, Continuous Integration, Data Cleansing, Extract Transform Load (ETL), Data Transformation, Python (Programming Language), SQL Databases, Data Processing, Scripting, Google Cloud, Feature Engineering, GitHub Copilot, Large Language Models, Snowflake, Prompt Engineering, Gitlab, Git, Pytest, Semi-structured Data, Machine Learning Operations, Terraform, Data Pipelines, Control M - **Published:** June 26, 2026 - **Apply:** https://www.dice.com/job-detail/3569c13a-8d61-41f5-987b-e4ae844bcaf8 ## About the Role * Must have skills: Data Engineer", "Snowflake SQL , "DBT", "Git", "Github Copilot", "AI . * SQL (Client dialect), dbt Core/Cloud basics. * Basic Python (scripting, pytest), with a strong understanding of its application in data pipelines and AI/ML. * Git / GitLab branching and MR workflows . * Familiarity with Terraform (reading/editing .tfvars). * Understanding of incremental models, source freshness, and dbt state: selectors. * Control-M scheduling concepts. * Client warehousing (sizing, multi-clustering). * dbt Cloud CLI, Poetry. * Github Copilot, Claude Code * Foundational understanding of how LLMs work, including concepts like tokenization, embeddings, and context windows. * Practical experience or strong theoretical knowledge of applying AI/LLMs to data processing tasks. * Ability to understand and work with unstructured and semi-structured data for AI applications. * Familiarity with prompt engineering principles for interacting with LLMs. * Strong analytical and problem-solving skills, with a proactive approach to learning new AI technologies. * US military Background is required. * Experience with vector databases and embedding pipelines. * Familiarity with MLOps principles and tools. * Exposure to cloud platforms (AWS, Google Cloud Platform, Azure). * Experience in building and optimizing data pipelines for AI workloads. * Understanding of AI model deployment and monitoring. ## Description * Maintain and enhance dbt models across staging (DL2) and mart (DL3) layers. * Monitor and triage dbt Cloud job failures, diagnosing root causes using run logs. * Support Control-M scheduling changes and DBT deployments. * Update dev_env.tfvars / Terraform configs for dbt Cloud environment and credential changes. * Write and maintain SQL models following SQLFluff lint standards (ruff, mypy for Python). * Manage source freshness and data quality test failures. * Contribute to GitLab MR reviews and follow CI/CD pipeline (Slim CI ? CD ? Prod deployment via migration tags). * Submit and track Client (SNOW) requests for Client access, service accounts, and infra changes. * Apply knowledge of LLM architecture and principles (e.g., tokenization, embeddings) to data preparation and feature engineering for AI/ML projects. * Develop and maintain robust Python code for ETL processes, with an understanding of how data is processed by LLMs. * Assist in evaluating and optimizing LLM outputs, considering factors like token limits and response relevance. * Document AI/ML data pipelines, including explanations of tokenization strategies and data transformations relevant to LLM processing. ## Related Videos - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [pytest: Simple, rapid and fun testing with Python](https://www.wearedevelopers.com/videos/213-pytest-simple-rapid-and-fun-testing-with-python) - [WeAreDevelopers LIVE - Modern DevOps for IoT Devices and More](https://www.wearedevelopers.com/videos/1805-wearedevelopers-live-modern-devops-for-iot-devices-and-more) - [Coffee with Developers - Maria Apazoglou](https://www.wearedevelopers.com/videos/1209-coffee-with-developers-maria-apazoglou) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)