> Markdown version of [/jobs/ext/2150538-data-and-knowledge-engineer](https://www.wearedevelopers.com/jobs/ext/2150538-data-and-knowledge-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data and Knowledge Engineer - **Company:** Robert Half - **Location:** Los Angeles, CA, United States - **Experience:** Expert - **Salary:** $200,000.0 - $300,000.0 - **Contract:** Temporary contract - **Skills:** Geographic Information Systems, Application Programming Interfaces (APIs), Artificial Intelligence, Data Analysis, Big Data, Information Systems, Databases, Customer Data Management, Data Validation, Data Deduplication, Information Engineering, Data Governance, Data Infrastructure, Data Integration, Data Mining, Relational Databases, File Systems, GIS Applications, Graph Database, Information Management, Job Scheduling, Python (Programming Language), Operational Databases, Standard Sql, Data Streaming, Systems Integration, Unstructured Data, Workflow Management Systems, Data Processing, Enterprise Software Applications, Data Ingestion, Large Language Models, Data Strategy, Knowledge Representation, Information Technology, Data Lineage, Data Analytics, Data Management, Data Pipelines, Automation Anywhere - **Published:** August 20, 2026 - **Apply:** https://dejobs.org/x/x/98A12E128644411FA4AA8615A530049A/job/ ## About the Role The Data & Knowledge Engineer will lead the onboarding, transformation, and governance of complex multimodal data into scalable data and knowledge platforms. This role focuses on integrating structured and unstructured data sources, designing reusable data pipelines, developing entity resolution frameworks, and enabling high-quality data for analytics, retrieval, AI workflows, and geospatial applications. The ideal candidate combines strong data engineering expertise with experience in knowledge graphs, data quality, and large-scale information management., * Bachelor's or Master's degree in Computer Science, Data Engineering, Statistics, Information Systems, or a related field, or equivalent practical experience. * 7+ years of experience building and operating production data platforms supporting structured and unstructured data workloads. * Strong Python and SQL expertise with modern software development best practices. * Experience with data ingestion, data integration, entity resolution, record linkage, master data management, or related disciplines. * Hands-on experience building batch and streaming data pipelines using modern data processing technologies. * Experience with workflow orchestration, job scheduling, and reliable reprocessing of large-scale data workloads. * Knowledge of knowledge graphs, vector databases, relational databases, and modern analytical data platforms. * Experience applying AI or large language models to document processing, data extraction, schema mapping, or data quality workflows. * Strong understanding of data governance, lineage, provenance, quality controls, and operational monitoring. * Ability to work with complex, messy, evolving, and ambiguous datasets while maintaining high data quality standards. * Excellent written and verbal communication skills., * Experience with OCR, document understanding, table extraction, multimodal processing, or intelligent document automation. * Experience working with geospatial data, spatial analytics, mapping technologies, or sensor-based datasets. * Knowledge of ontology development, graph data modeling, knowledge representation, or semantic technologies. * Experience with active learning, human-in-the-loop review processes, confidence scoring, or data quality automation. * Familiarity with data lineage, privacy-focused architectures, data contracts, and governance frameworks. * Experience supporting secure, air-gapped, edge, or highly regulated environments. * Experience working directly with enterprise, government, or large-scale organizational data onboarding initiatives. * Background in defense, intelligence, logistics, public sector, public safety, or other complex data-intensive industries. * Experience mentoring engineers and establishing data engineering standards and best practices. ## Description * Design and develop scalable data ingestion pipelines for structured, unstructured, geospatial, and sensor-based data sources. * Build and maintain batch and streaming data processing systems across cloud, on-premises, and disconnected environments. * Develop integrations with APIs, databases, file systems, enterprise applications, and external data sources. * Design schema mapping, normalization, and transformation processes that support diverse customer data models. * Implement entity resolution, record linkage, deduplication, and data matching capabilities across multiple sources. * Preserve data lineage, provenance, auditing, and traceability throughout the data lifecycle. * Create data validation, monitoring, replay, and exception-handling processes for complex data environments. * Develop workflows for managing ambiguous records, conflicting information, and data quality issues. * Define and measure data quality metrics, onboarding effectiveness, and operational performance indicators. * Support knowledge graph, retrieval, AI, and analytics capabilities through high-quality governed datasets. * Partner with engineering and stakeholder teams to transform recurring onboarding requirements into reusable platform capabilities. * Contribute to platform architecture, engineering standards, and long-term data strategy initiatives. Additional Details * Fully onsite 5 days a week * Full-time exempt position * Hands-on engineering role with substantial ownership and technical influence * Opportunity to work on large-scale data, knowledge graph, and AI-driven initiatives * Staff-level candidates may provide architectural leadership, mentorship, and engineering guidance * Candidates must be authorized to work in the United States and satisfy applicable regulatory employment requirements ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Data Science on Software Data](https://www.wearedevelopers.com/videos/162-data-science-on-software-data) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Kubernetes and Microservices with Multi-Model Databases](https://www.wearedevelopers.com/videos/382-kubernetes-and-microservices-with-multi-model-databases) - [Fault Tolerance and Consistency at Scale: Harnessing the Power of Distributed SQL Databases](https://www.wearedevelopers.com/videos/1146-fault-tolerance-and-consistency-at-scale-harnessing-the-power-of-distributed-sql-databases) - [Data Mining Accessibility](https://www.wearedevelopers.com/videos/802-data-mining-accessibility) ## Related Articles - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market)