Data Engineer

Collabera
Houston, TX, United States
9 days ago
Apply on www.collabera.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
2 years minimum
Compensation
$135,000.0 - $165,000.0
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Automation of Tests Encodings Data Architecture Data Validation Data Cleansing Information Engineering Data Governance Data Integration Data Integrity Data Systems
+22 more
Database Queries Document-Oriented Databases Fault Tolerance Python (Programming Language) Machine Learning Regression Testing Mobile Analytics DataOps Search Technologies SQL Databases Systems Integration Data Server Interface Enterprise Software Applications Data Classification Retrieval-Augmented Generation Large Language Models Usage Tracking Pyspark Data Management Network Server Data Pipelines Databricks

Job description

o Design, build, and maintain scalable, self-healing data pipelines using Databricks, Python, PySpark, and SQL. o Develop data pipelines across Bronze, Silver, and Gold/Certified Gold layers, ensuring data quality and reliability. o Ingest, transform, and serve data from dozens of source systems, including enterprise applications, financial systems, IoT, web/mobile analytics, and third-party platforms. o Build error recovery and quarantine workflows to isolate failed records while allowing valid data to continue through the pipeline. o Implement schema validation, data quality checks, anomaly detection, automated testing, regression testing, and data observability. o Develop and maintain data models with proper data grain, keys, referential integrity, lineage, and governance. o Build infrastructure that supports AI/ML workloads, including feature stores, embedding pipelines, vector search, and real-time serving layers. o Build and maintain MCP server integrations that expose enterprise data to LLM-powered tools and AI agents. o Develop integrations that allow AI applications to connect to, query, and retrieve data through MCP servers. o Design and implement RAG (Retrieval-Augmented Generation) architectures and integrate LLM-powered solutions with enterprise data. o Work with vector databases such as Pinecone, Weaviate, or similar technologies. o Support AI model training, evaluation, deployment, monitoring, and productionization in partnership with Data Science and Product teams. o Evaluate AI-powered data engineering and data quality tools, including solutions for automated schema detection, cataloging, completeness checks, and data validation. o Use AI-assisted testing and evaluation approaches to validate data and AI outputs against business requirements. o Develop APIs and data interfaces that enable AI products and internal applications to query and interact with data in real time. o Implement data governance practices covering access controls, PII handling, data classification, compliance, and appropriate AI data usage. o Build monitoring, alerting, SLA tracking, and data freshness capabilities into data platforms. o Document data models, pipeline architectures, AI integrations, and reusable engineering patterns.

Requirements

o 5+ years of experience in Data Engineering, working with data from multiple enterprise source systems. o Strong hands-on experience with Databricks. o Strong Python and PySpark development experience. o Strong SQL skills. o 2-3+ years of hands-on experience with MCP Servers / Model Context Protocol. o Hands-on experience building and implementing RAG models/architectures. o Experience connecting to, querying, and integrating data through MCP servers. o Strong experience building self-healing or fault-tolerant data pipelines and error recovery workflows. o Experience with schema validation and data quality frameworks. o Strong understanding of Bronze/Silver/Gold data architecture. o Experience with LLM integration patterns, AI agents, tool-use frameworks, and AI-enabled data solutions. o Experience with vector databases such as Pinecone, Weaviate, or equivalent. o Experience building data pipelines and infrastructure suitable for AI/ML workloads. o Understanding of data governance, lineage, monitoring, observability, and data quality.

Benefits & conditions

This is a direct hire opportunity. The selected candidate will be employed directly by our client. All compensation and benefits, including but not limited to medical insurance, retirement plans, paid time off, and other perks, will be provided by the client in accordance with their internal policies and subject to applicable laws and eligibility requirements.

Job Requirement o MCP o RAG o Pyspark

Reach Out to a Recruiter o Recruiter

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collabera.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:17 min

Optimizing character encoding with Kim variable byte encoding

Douglas Crockford Douglas Crockford · World Congress 2024

2:00 min

Separating dataset creation from low-level software implementation steps

Jan Zawadzki · World Congress 2022

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

4:12 min

Distilling cross-encoder models into smaller efficient sentence embedding models

Marek Suppa · LIVE

Videos

See all

Related articles

See all