> Markdown version of [/jobs/ext/2642262-data-engineer](https://www.wearedevelopers.com/jobs/ext/2642262-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** Warwick Investment Group, LLC - **Location:** Oklahoma City, OK, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Query Performance, Application Programming Interfaces (APIs), Artificial Intelligence, Airflow, Microsoft Azure, Databases, Continuous Integration, Data Architecture, Data Validation, Data Governance, Data Infrastructure, Extract Transform Load (ETL), Data Warehousing, Database Development, Digital Assets, Document-Oriented Databases, Supervisory Control and Data Acquisition (SCADA), Python (Programming Language), Machine Learning, Query Optimization, Power BI, Software Tools, Cloud Services, SQL Databases, Data Streaming, Scripting, Data Ingestion, Sql Optimization, GitHub Copilot, System Availability, Large Language Models, Snowflake, Multi-Agent Systems, Change Data Capture, Git, Containerization, Data Lineage, Apache Kafka, Spotfire, Software Version Control, Data Pipelines, Docker - **Published:** August 2, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=5e9fd3c8ca67917e ## About the Role * Oil and Gas Experience: 5+ years in oil and gas data environments, with familiarity across production, land, SCADA, and well data domains. * SQL and Data Modeling: Advanced SQL proficiency (CTEs, window functions, query optimization). Experience with dimensional and normalized modeling approaches. * Pipeline and Orchestration: Experience building and maintaining ELT/ETL pipelines with Coalesce, dbt, Airflow, Prefect, or similar tools. * Python: Proficient in Python for data manipulation, pipeline scripting, and automation tasks. * Cloud Data Platforms: Experience with Snowflake & Azure for warehousing, storage, and compute. * Version Control and CI/CD: Familiarity with Git, Azure DevOps, or similar systems. Experience with CI/CD for data pipeline deployments. * Data Governance: Understanding of data quality frameworks, lineage tracking, access control, and compliance requirements. * Documentation: Capable of producing clear technical documentation for pipelines, schemas, and processes for both technical and non-technical audiences. * Containerization: Proficiency with Docker for containerizing data workloads and pipeline components, with familiarity deploying containers to cloud services. DESIRED SKILLS * AI-Assisted Development: Experience using AI coding tools (e.g., Claude Code, GitHub Copilot) to accelerate SQL development, dbt workflow creation, and data pipeline work. * AI Data Architecture: Familiarity with RAG patterns, vector stores, embeddings, or building data interfaces for LLM and agent consumption. * Agent and MCP Workflows: Experience building or orchestrating multi-agent AI systems, MCP servers, or tool-use interfaces that expose data to AI systems. * AI-Powered Documentation: Using AI tools to auto-generate and maintain data dictionaries, lineage documentation, and schema descriptions that stay current with the codebase. * BI Tool Familiarity: Working knowledge of Power BI or Spotfire to collaborate effectively with the BI team. * Streaming and CDC: Experience with change data capture, event streaming (Kafka, Azure Event Hubs), or real-time data pipelines. ## Description The Data Engineer at Warwick Energy owns the infrastructure that moves data from source systems into a trusted, well-modeled warehouse, and prepares that data for consumption by both human analysts and AI systems. You will design and maintain ingestion pipelines, orchestrate ELT workflows, enforce data quality, and build the semantic and structural layers that let BI tools, machine learning models, and AI agents draw on a single source of truth. A critical dimension of this role is forward-looking: as the organization moves toward AI-driven decision-making, you will shape data assets so they are discoverable, well-documented, and structured for retrieval-augmented generation (RAG), agent workflows, and automated analytics. You will partner with data scientists, BI analysts, and business stakeholders to ensure the data platform scales with both traditional reporting needs and emerging AI use cases., * Pipeline Development and Orchestration: Build, monitor, and maintain data ingestion pipelines from source systems (APIs, databases, flat files, SCADA/IoT) into the data warehouse. Orchestrate ELT workflows using tools like Coalesce, dbt, Airflow, or Prefect, with version control and CI/CD practices. * Data Modeling and Warehouse Architecture: Design scalable dimensional and normalized data models. Own the warehouse layer structure (raw, staging, marts) and ensure models support both BI reporting and AI/ML consumption patterns. * Data Quality and Governance: Implement data quality checks, monitoring, and alerting across pipelines. Enforce governance standards including lineage tracking, access control, and documentation to maintain trust in data assets. * AI-Ready Data Architecture: Structure and document data assets so they are consumable by LLMs, RAG pipelines, and AI agents. Design metadata layers, semantic descriptions, and context-rich schemas that allow AI systems to discover and reason over organizational data. * AI-Accelerated Engineering: Use AI coding tools (Claude Code, Copilot) and agent workflows to accelerate pipeline development, automate documentation, generate and validate SQL transformations, and build MCP servers or similar interfaces that expose data to AI systems. * Infrastructure and Platform Reliability: Manage cloud data infrastructure (Snowflake, Azure). Monitor pipeline health, optimize query performance, and maintain SLAs for data freshness and availability. * Documentation and Collaboration: Produce clear technical documentation for pipelines, data models, and ELT processes. Partner with BI analysts, data scientists, and business teams to align data infrastructure with analytical and AI-driven objectives. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again)