TELECOMMUTE Principal Data Architect (REMOTE)

AITA Consulting Services Inc.
United States
1 day ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source

Tech stack

Training Data Artificial Intelligence Airflow Amazon Web Services Amazon S3 Computing Platforms Microsoft Azure Big Data Business Software Information Systems Data Architecture Information Engineering
+57 more
Data Governance Extract Transform Load (ETL) Data Security Data Vault Modeling Software Debugging Github Graph Database Python (Programming Language) Machine Learning Metadata Neo4j Query Optimization Power BI OpenAI Cloud Services Data Mesh Standard Sql SQL Databases Tableau (Software) Datadog Pinecone Feature Store Cloud Platform System Azure Data Factory LangChain Great Expectations (Foster Youth College-readiness and Support Program) Retrieval-Augmented Generation Large Language Models Snowflake Grafana Apache Spark Llamaindex Git Fastapi Data Layers Data Lakes Pyspark Core Data Kubernetes Information Technology Collibra Apache Kafka Apache Nifi Spark Streaming Data Management Machine Learning Operations FAISS Claude Cloudwatch Model Context Protocol Terraform Looker Analytics OpenSearch Docker Jenkins ChromaDB Databricks

Job description

We are hiring a Principal Data Architect - a hands-on, senior individual contributor who will design, build, govern, and evolve the enterprise data platform that serves as the single source of truth for the organization. You will own the architecture end to end: how data is modeled, ingested, transformed, stored, governed, and served to every consumer that depends on it - from executive reporting and operational analytics to business applications and, increasingly, ML and AI systems. This is a data architecture role first. The large majority of your time goes to data modeling, platform design, pipeline engineering, data quality, and governance. A smaller but growing part of it goes to making that foundation AI-ready - so that ML models and AI applications consume governed, documented, trustworthy data instead of building their own shadow pipelines., * You think in domains and contracts: every design decision considers all current and future consumers, not just the one request in front of you.

  • You are as comfortable in a governance and stewardship session as you are debugging a skewed Spark join late at night.
  • You believe the hard part of data work is agreement, not technology - and you do the work of getting teams to a single definition.
  • You close the loop: nothing is shipped until it has tests, monitoring, documentation, and a named owner.
  • You are pragmatic about AI - enthusiastic about enabling it, unwilling to let it bypass governance., * Core data platform: Snowflake · Databricks · Delta Lake · PySpark · Spark Structured Streaming · Kafka · Apache NiFi · Airflow · dbt · SQL · Python · AWS (S3, Glue, Redshift, Kinesis, EKS, Lambda) · Azure · Terraform · Docker · Kubernetes · GitHub Actions · Jenkins · Grafana / CloudWatch
  • Governance & quality: Unity Catalog · Collibra / Alation / Purview · Great Expectations · MDM platforms · column-level lineage tooling · Power BI / Tableau / Looker semantic layers
  • AI enablement: MLflow · feature stores · OpenSearch / Pinecone / FAISS / ChromaDB · Neo4j · LangChain / LlamaIndex · AWS Bedrock · Claude · MCP

Requirements

Must-Have Experience

  • 15+ years of hands-on data engineering and data architecture experience, with a track record of owning enterprise-scale platforms end to end.
  • Deep data modeling expertise - dimensional, normalized, and Data Vault - with real judgment about grain, conformance, and when each pattern applies.
  • Proven experience architecting lakehouse and/or data mesh platforms: Databricks, Delta Lake, PySpark, Snowflake, Kafka, Spark Structured Streaming, and cloud-native data services (AWS, Azure).
  • Strong command of SQL and performance engineering: query optimization, partitioning and clustering, and cost management at scale.
  • Demonstrated ownership of data governance, data quality, master data, metadata, and lineage programs - not just the tooling, but the operating model.
  • Experience with data security and compliance in regulated industries (financial services, payments, cybersecurity, healthcare).
  • Experience migrating legacy on-prem warehouses and ETL estates to cloud-native platforms.
  • Working familiarity with the data requirements of ML and AI systems - feature pipelines, training data curation, and retrieval/vector data layers - enough to architect for them credibly and partner well with AI teams.

Technical Skills

  • Expert: SQL, Python, PySpark, data modeling, Databricks, Delta Lake, Snowflake, Kafka, Spark Structured Streaming, AWS (S3, Glue, Redshift, Kinesis, EKS, Lambda), Airflow, Terraform, Git-based CI/CD.
  • Strong: Azure data services, dbt, data quality frameworks (Great Expectations or equivalent), catalog and lineage platforms (Unity Catalog, Collibra, Alation, Purview), MDM tooling, BI semantic layers (Power BI, Tableau, Looker), Docker, Kubernetes, FastAPI, observability tooling (Grafana, CloudWatch).
  • Familiar: MLflow, feature stores, vector databases (OpenSearch, Pinecone, FAISS, ChromaDB), knowledge graphs (Neo4j), LLM APIs and RAG frameworks (LangChain, LlamaIndex, AWS Bedrock, Claude, OpenAI)., * BS / MS in Computer Science, Information Systems, Data Engineering, or a related quantitative field.
  • Prior experience at a global financial institution (payments, risk, AML, compliance) or large enterprise SaaS, operating under strict regulatory oversight.
  • Experience in presales solutioning for large data programs: RFP/RFI responses, SOW shaping, effort estimation, and CXO-level solutioning.
  • Certifications such as Databricks Data Engineer Professional, AWS Certified Data Engineer or Solutions Architect, or Snowflake SnowPro Advanced Architect.
  • Familiarity with emerging standards for AI and agent data access (e.g., MCP - Model Context Protocol) and agent entitlement models.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Loading talks and stories from around this role…