> Markdown version of [/jobs/ext/1712490-ai-data-architect](https://www.wearedevelopers.com/jobs/ext/1712490-ai-data-architect). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Data Architect - **Company:** 3 Pillar Global - **Location:** United States (Remote available) - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Amazon S3, Audit Trail, Microsoft Azure, Batch Processing, Cloud Engineering, Cloud Storage, Code Review, Continuous Integration, Data Architecture, Information Engineering, Data Infrastructure, Extract Transform Load (ETL), Data Warehousing, Software Design Patterns, Github, Graph Database, Identity and Access Management, Python (Programming Language), Machine Learning, Meta-Data Management, Neo4j, Software Product Management, SQL Databases, Data Streaming, Datadog, Large Language Models, Snowflake, Grafana, Fastapi, Event Driven Architecture, Data Lakes, Pyspark, Kubernetes, Data Lineage, Low Latency, HuggingFace, Apache Kafka, Spark Streaming, Data Management, Machine Learning Operations, Virtual Agents, Cloudwatch, Terraform, Data Pipelines, Docker, Databricks - **Published:** July 25, 2026 - **Apply:** https://jobs.lever.co/3pillarglobal/d2ded0cc-eb2c-4185-9347-62d5c9f402bd/apply ## About the Role Architect and own the enterprise AI data platform - the unified, governed layer that ingests, transforms, stores, and serves all data consumed by AI systems across the organisation. Design multi-domain data models (lakehouse, data mesh, event-driven) that are structured from day one to serve AI workloads: clean lineage, versioned schemas, well-documented contracts, and low-latency serving APIs. Strong exposure to different Data architectures, data lake & data warehouse Define tools & technologies to develop automated data pipelines, write ETL processes, develop dashboard & report and create insights Responsibilities Technical Skills Primary Skills: Python, SQL, Snowflake/Databricks, AWS (S3, Glue, EKS, Bedrock, Kinesis, Redshift), Docker, Kubernetes, Terraform, GitHub Actions, LangChain, LlamaIndex, LLM APIs (OpenAI, AWS Bedrock, Claude, HuggingFace), (Pinecone, FAISS, ChromaDB, OpenSearch), knowledge graphs (Neo4j). Secondary Skills: MLflow, FastAPI, CI/CD pipelines, observability tooling (CloudWatch, Grafana, or equivalent), data lineage and metadata management platforms. * 15+ years of hands-on data engineering and architecture experience, alongside building production AI/ML and LLM-era data infrastructure. * Strong Experience with either Databricks or Snowflake; experience with both is desirable. * Strong data architecture patterns & principles, ability to design secure & scalable data lakes, data warehouse, data hubs, and other event-driven architectures * Expertise in designing and writing ETL processes in Python / Java / Scala * Own the full data stack: real-time streaming (Kafka, Spark Structured Streaming), batch processing (Databricks, PySpark, Delta Lake), cloud storage and compute (AWS, Azure), and data quality /metadata management. * Drive modernisation of legacy pipelines (on-prem ETL, batch DWH) to cloud-native, AI-ready architectures with measurable improvements in cost, latency, and delivery velocity. * Proven experience designing enterprise-scale AI data platforms that serve multiple AI consumers -not just one application or pipeline. * Hands-on experience with vector stores, semantic models, knowledge graphs, and retrieval infrastructure in production environments. * Working knowledge of LLMOps: model serving pipelines, MLflow, CI/CD for AI, automated evaluation, and production monitoring. AI Experience RAG, Vector & Retrieval Infrastructure Design the retrieval infrastructure that powers RAG-based AI applications: embedding pipelines, vector stores (Pinecone, FAISS, ChromaDB, OpenSearch), chunking strategies, and hybrid retrieval layers combining semantic search with structured queries. ## Description This platform will serve as the foundational nervous system for conversational AI assistants, dashboard intelligence, autonomous AI agents, RAG-powered applications, predictive ML models, and any AI product we build today or in the future. The resource will architect the system, drive implementation, own the data contracts that agents and AI applications depend on, enforce security and access governance for both human and agent consumers, and continuously monitor and improve the accuracy and reliability of AI outputs that flow from this platform., Own the observability stack for AI agent behaviour: instrument agents to capture inputs, retrieved context, tool calls, reasoning traces, and outputs - creating a complete audit trail of every agentic action driven by platform data. Design and operate evaluation frameworks that continuously measure AI output quality: factual accuracy, context faithfulness, retrieval relevance, hallucination rates, and task completion success- across all AI consumers of the platform. Architecture Standards & Engineering Enablement Define and maintain the reference architecture for the AI data platform - documenting design patterns, data contracts, integration standards, and decision records (ADRs) that all engineering teams follow. Establish data engineering standards: pipeline testing frameworks, code review practices, CI/CD automation, infrastructure-as-code (Terraform), reusable component libraries, and observability instrumentation. ## Related Videos - [Putting the Graph In GraphQL With The Neo4j GraphQL Library](https://www.wearedevelopers.com/videos/257-putting-the-graph-in-graphql-with-the-neo4j-graphql-library) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Cyber Sleuth: Finding Hidden Connections in Cyber Data](https://www.wearedevelopers.com/videos/893-cyber-sleuth-finding-hidden-connections-in-cyber-data) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) ## Related Articles - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)