> Markdown version of [/jobs/ext/2935069-data-engineer](https://www.wearedevelopers.com/jobs/ext/2935069-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** Collabera - **Location:** Houston, TX, United States (Remote available) - **Experience:** Expert - **Salary:** $135,000.0 - $165,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Automation of Tests, Encodings, Data Architecture, Data Validation, Data Cleansing, Information Engineering, Data Governance, Data Integration, Data Integrity, Data Systems, Database Queries, Document-Oriented Databases, Fault Tolerance, Python (Programming Language), Machine Learning, Regression Testing, Mobile Analytics, DataOps, Search Technologies, SQL Databases, Systems Integration, Data Server Interface, Enterprise Software Applications, Data Classification, Retrieval-Augmented Generation, Large Language Models, Usage Tracking, Pyspark, Data Management, Network Server, Data Pipelines, Databricks - **Published:** September 16, 2026 - **Apply:** https://www.collabera.com/submit-resume/ ## About the Role o 5+ years of experience in Data Engineering, working with data from multiple enterprise source systems. o Strong hands-on experience with Databricks. o Strong Python and PySpark development experience. o Strong SQL skills. o 2-3+ years of hands-on experience with MCP Servers / Model Context Protocol. o Hands-on experience building and implementing RAG models/architectures. o Experience connecting to, querying, and integrating data through MCP servers. o Strong experience building self-healing or fault-tolerant data pipelines and error recovery workflows. o Experience with schema validation and data quality frameworks. o Strong understanding of Bronze/Silver/Gold data architecture. o Experience with LLM integration patterns, AI agents, tool-use frameworks, and AI-enabled data solutions. o Experience with vector databases such as Pinecone, Weaviate, or equivalent. o Experience building data pipelines and infrastructure suitable for AI/ML workloads. o Understanding of data governance, lineage, monitoring, observability, and data quality. ## Description o Design, build, and maintain scalable, self-healing data pipelines using Databricks, Python, PySpark, and SQL. o Develop data pipelines across Bronze, Silver, and Gold/Certified Gold layers, ensuring data quality and reliability. o Ingest, transform, and serve data from dozens of source systems, including enterprise applications, financial systems, IoT, web/mobile analytics, and third-party platforms. o Build error recovery and quarantine workflows to isolate failed records while allowing valid data to continue through the pipeline. o Implement schema validation, data quality checks, anomaly detection, automated testing, regression testing, and data observability. o Develop and maintain data models with proper data grain, keys, referential integrity, lineage, and governance. o Build infrastructure that supports AI/ML workloads, including feature stores, embedding pipelines, vector search, and real-time serving layers. o Build and maintain MCP server integrations that expose enterprise data to LLM-powered tools and AI agents. o Develop integrations that allow AI applications to connect to, query, and retrieve data through MCP servers. o Design and implement RAG (Retrieval-Augmented Generation) architectures and integrate LLM-powered solutions with enterprise data. o Work with vector databases such as Pinecone, Weaviate, or similar technologies. o Support AI model training, evaluation, deployment, monitoring, and productionization in partnership with Data Science and Product teams. o Evaluate AI-powered data engineering and data quality tools, including solutions for automated schema detection, cataloging, completeness checks, and data validation. o Use AI-assisted testing and evaluation approaches to validate data and AI outputs against business requirements. o Develop APIs and data interfaces that enable AI products and internal applications to query and interact with data in real time. o Implement data governance practices covering access controls, PII handling, data classification, compliance, and appropriate AI data usage. o Build monitoring, alerting, SLA tracking, and data freshness capabilities into data platforms. o Document data models, pipeline architectures, AI integrations, and reusable engineering patterns. ## Related Videos - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [A Brief History of Data Storage](https://www.wearedevelopers.com/videos/974-a-brief-history-of-data-storage) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) ## Related Articles - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)