> Markdown version of [/jobs/ext/2661261-principal-data-architect-and-manager](https://www.wearedevelopers.com/jobs/ext/2661261-principal-data-architect-and-manager). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Data Architect and Manager - **Company:** Apple Inc. - **Location:** United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Amazon S3, Apache HTTP Server, Big Data, Compilers, Computer Programming, Continuous Integration, Data Architecture, Information Engineering, Dimensional Modeling, Graph Database, Python (Programming Language), Machine Learning, Neo4j, Software Architecture, Cloud Services, Search Technologies, Data Streaming, Datadog, Pulumi, Large Language Models, Data Lakes, Information Technology, Collibra, Low Latency, Apache Flink, ONNX (Open Neural Network Exchange) Format, AWS Glue, Apache Kafka, Spark Streaming, Data Management, TensorRT, Terraform, Stream Processing, Data Pipelines - **Published:** August 2, 2026 - **Apply:** https://www.dice.com/job-detail/85697900-6831-454c-bd96-04e706b521c2 ## About the Role MS Degree in Computer Science or related degree and 12+ years of experience in Data Architecture, Data Engineering, or Platform Engineering, with at least 5 years operating in a Principal, Staff, or Lead Manager capacity. Proven experience leading and managing engineers including hiring, performance management, and technical mentorship of senior ICs and managers. Track record of shipping petabyte-scale, low-latency data platforms in production and operating them under real-world load. Deep cloud expertise: expert-level proficiency with cloud object storage (e.g., AWS S3) and its architectural nuances for massive data lakes and lake-houses. Experience architecting systems for entity resolution, conflation, or knowledge-graph construction at scale - ideally involving billions of frequently updated entities. Experience designing pipelines that process multimodal data (structured, text, image) and integrate ML model inference including LLMs and embedding models: for enrichment and transformation. Familiarity with LLM/model-serving infrastructure trade-offs (inference runtimes, GPU-backed serving) to inform architectural decisions Streaming expertise: deep, hands-on knowledge of Apache Kafka (or comparable brokers like Kinesis) and complex stream processing (Spark Structured Streaming, Flink, or similar). Data modeling: exceptional ability to design logical and physical data models for large-scale ingest, retrieval, and analytical consumption - including dimensional modeling and lakehouse patterns. Experience defining SLAs, quality metrics, and observability standards for large-scale data platforms, with hands-on use of monitoring/alerting tooling (e.g., PrometheGrafana, Datadog, or OpenTelemetry-based tracing). Programming: command of at least one modern data-pipeline language (Scala, Java, or Python) and strong software engineering fundamentals. Cloud services integration: proven experience wiring together event notifications, queuing, orchestration, and compute services into resilient production pipelines. Experience with vector search technologies (e.g., Pinecone, Milvus) and storing/serving embeddings (e.g., pgvector, Milvus, FAISS) Excellent written and verbal communication; proven ability to align engineers, partner teams, and senior leadership from multiple lines of business around a shared technical direction, with experience bringing a consumer-oriented product from inception to production. Preferred Qualifications Experience with embedding storage and retrieval (e.g., pgvector, Milvus, FAISS) and with graph databases (e.g., TigerGraph, Neo4j). Experience deploying, serving, and optimizing LLMs or ML models directly in the production, inference runtimes/compilers (ONNX Runtime, TensorRT/TensorRT-LLM), and serving frameworks (Triton, vLLM, TorchServe or similar). Experience tuning batching, KV-cache, and GPU utilization for low-latency, high-throughput real-time inference in a data pipeline Experience with data governance tools (e.g., Apache Atlas, AWS Glue Catalog, DataHub). Familiarity with Infrastructure as Code (Terraform, Pulumi) and modern CI/CD practice. Experience designing systems that handle petabytes of unstructured media data. Working knowledge of data privacy regulations and best practices for incorporating safety and compliance, and a demonstrated instinct for building privacy-preserving systems. ## Description We are looking for a Principal Data Architect and Manager to serve as both the senior technical authority and the people leader for our data platform., As the Principal Data Architect and Manager on our team, you will serve as both the senior technical authority and the people leader for our data platform. You'll define and own the end-to-end architecture of a real-time, petabyte-scale data backbone: from ingestion through a multi-layered lakehouse to normalized serving layers that power downstream search, ranking, and on-device experiences. You'll also build, grow, and lead the team of data engineers who bring that architecture to life. This is a hands-on principal role with multiple facets: you set the technical vision, personally shape the hardest architectural decisions, drive the roadmap through to production, and manage, mentor, and grow the engineers executing against it. Your leverage comes equally from what you design and from the team you build. ## Related Videos - [Putting the Graph In GraphQL With The Neo4j GraphQL Library](https://www.wearedevelopers.com/videos/257-putting-the-graph-in-graphql-with-the-neo4j-graphql-library) - [Why segmenting your infrastructure into tiers makes your infrastructure design better](https://www.wearedevelopers.com/videos/1960-why-segmenting-your-infrastructure-into-tiers-makes-your-infrastructure-design-better) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Cyber Sleuth: Finding Hidden Connections in Cyber Data](https://www.wearedevelopers.com/videos/893-cyber-sleuth-finding-hidden-connections-in-cyber-data) - [Blueprints for Success: Steering a Global Data & AI Architecture](https://www.wearedevelopers.com/videos/1577-blueprints-for-success-steering-a-global-data-ai-architecture) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence)