> Markdown version of [/jobs/ext/2932424-ai-data-knowledge-engineer](https://www.wearedevelopers.com/jobs/ext/2932424-ai-data-knowledge-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Data & Knowledge Engineer - **Company:** Apple Inc. - **Location:** Cupertino, CA, United States - **Experience:** Expert - **Salary:** $149,200.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Artificial Intelligence, Airflow, Amazon Web Services, Data Analysis, Application Frameworks, Systems Engineering, Microsoft Azure, Cloud Database, Software Documentation, Computer Programming, Databases, Data Validation, Data Dictionary, Data Governance, Extract Transform Load (ETL), Data Transformation, Data Systems, Programming Tools, Distributed Computing Environment, Distributed Data Store, R (Programming Language), Graph Database, Apache Hadoop, JSON, Python (Programming Language), Machine Learning, Metadata, NoSQL, Parsing, Search Technologies, SQL Databases, Unstructured Data, Management of Software Versions, Web Standards, Parquet, Google Cloud, Chatbots, Large Language Models, Snowflake, Grafana, Apache Spark, Indexer, Git, Data Layers, Build Management, Semi-structured Data, Collibra, Virtual Agents, Artificial Intelligence Markup Language (AIML), Software Version Control, Data Pipelines, Amazon Redshift, Databricks - **Published:** September 16, 2026 - **Apply:** https://www.jobmonkeyjobs.com/career/28021940/Ai-Data-Knowledge-Engineer-California-Cupertino-7413 ## About the Role Experience designing and building knowledge layers for AI systems, including knowledge graphs, RAG pipelines, and vector databases to ground LLM-driven applications in accurate, structured, unstructured and retrievable enterprise knowledge. Experience modeling enterprise knowledge and metadata within semantic layers to represent business entities, attributes, and their relationships. 5+ years of experience in designing, building, and maintaining scalable data solutions for large-scale analytics. Proficiency in SQL and development experience with cloud database environments like Snowflake, Redshift, Databricks. Proficiency in programming languages like Python, Java, R and open-source frameworks for distributed processing like Hadoop and Spark. Experience building data pipelines to ingest, transform, and continuously synchronize structured and unstructured enterprise data from multiple sources. Hands-on experience using development tools in a modern cloud data stack for code management, versioning using Git, CI/CD tools, automation and orchestration using Apache Airflow or others and monitoring & alerting. Experience with Cloud platforms AWS, Azure or Google Cloud. Preferred Qualifications Experience architecting and developing data pipelines through ETL tools, API integration with on-premise and cloud-based sources. Experience building ontology-based semantic layer including a business ontology of sales concepts, a technical ontology of data sources and schemas, and execution traces that provide feedback for continuous improvement. Strong understanding of LLM evaluation and AI quality tooling, retrieval metrics, and observability to improve application reliability. Experience working with unstructured and Semi-structured data sets (e.g., JSON, Parquet, PDF, text, images, audio, video) Experience with data governance and observability tools; for example DataHub, Collibra Experience articulating and translating business questions into data solutions and proven ability to lead development projects from start to finish. Broad knowledge of web standards relating to REST, HTTP, JSON, etc. Experience with data labeling and annotation tools and processes. Familiarity with AI/ML model development lifecycle and data needs for training and deployment. Ability to balance competing priorities, long-term projects, and ad hoc requirements. Ability to work in a fast-paced, dynamic, constantly evolving business environment. ## Description As an AI Data & Knowledge Engineer, you will develop infrastructure, systems, services, and tools for automating sales processes. We're looking for an exceptional engineer that lives at the intersection of development, operations, data, and systems engineering to build solutions for large-scale continuous data transformation and delivery. This role will specifically focus on building and maintaining data pipelines for both structured and unstructured data, enabling the development and deployment of AIML models. Responsibilities Responsible for the development and design of data pipelines and data knowledge layers for agentic AI applications. Design and implement data models for a semantic layer that integrates analytics data from multiple sources in an efficient and effective manner. Design and build scalable data and knowledge layers that power chatbots and other agentic applications. Build RAG-ready data pipelines and knowledge layers encompassing document ingestion, parsing, metadata tagging, embeddings, indexing with vector search. Design scalable architecture for semantic and hybrid search, knowledge graphs to enable contextually accurate text-to-SQL generation. Build mechanisms for incremental synchronization of data to knowledge updates so agent responses are current and reliable. Designing and operating distributed data systems - from SQL/NoSQL databases, Vector search, and orchestration. Collaborate with Analytics and Data Science teams to translate business requirements into reliable, actionable knowledge layers that support AI agent development and deliver targeted business outcomes. Collaborate with internal business partners, internal technology resources (database, system, networking), external vendors, and partners. Play an active role in the development and maintenance of user documentation, including data models, mapping rules, and data dictionaries. Ensure data quality and accuracy by developing data validation and reconciliation processes. Build and maintain data pipelines for ingesting, processing, and transforming unstructured data sources, such as customer feedback, social media data, or sales call recordings. Develop data quality monitoring and validation processes specifically for AIML datasets, including identifying and addressing data bias. Work with data scientists to understand data requirements for AIML model training and deployment, ensuring data is available in the appropriate format and quality. Implement data governance policies and procedures to ensure the responsible and ethical use of data in AIML applications. ## Related Videos - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Tips and Tricks for Working with JSON](https://www.wearedevelopers.com/videos/1229-tips-and-tricks-for-working-with-json) - [Leveraging Real time data in FSIs](https://www.wearedevelopers.com/videos/806-leveraging-real-time-data-in-fsis) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) - [Developer Experience, Platform Engineering and AI powered Apps](https://www.wearedevelopers.com/videos/990-developer-experience-platform-engineering-and-ai-powered-apps) ## Related Articles - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models)