Data Architect (Hands on)

Stratus Technologies
United States
3 days ago

Role details

Contract type
Permanent contract
Employment type
Part-time (≤ 32 hours)
Experience level
Expert
Experience required
3 years minimum
Compensation
$160,000.0 - $200,000.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Data Analysis Microsoft Azure Data Architecture Information Engineering Data Governance Relational Databases Cursor (Graphical User Interface Elements) Shard (Database Architecture) Disaster Recovery MongoDB
+14 more
Operational Databases Query Optimization Regression Testing Cloud Services Search Technologies Azure Data Factory Snowflake Generative AI Indexer Sap Business Objects Data Layers Debezium Data Pipelines Databricks

Job description

Reposted 15 Hours Ago Remote Hiring Remotely in United States Senior level Remote Hiring Remotely in United States Senior level Own and implement the canonical data model and governance for a multi-tenant MEP SaaS platform. Architect data for AI/ML readiness (RAG, embeddings, vector search), design polyglot persistence and lake/lakehouse pipelines, lead staged modernization and migrations, produce hands-on prototypes and in-repo guardrails, and partner with platform, DB engineering, and ML teams to ensure data quality, lineage, and observability. The summary above was generated by AI

Stratus, deriving from the Latin term meaning ‘layer’, offers an advanced set of MEP specific solutions that seamlessly layer across a contractor’s entire workflow from design to fabrication to installation. Our team of seasoned industry experts, skilled technology leaders, innovators, and entrepreneurs understands that fabrication does not occur in isolation, and increasingly, it may not happen within your own fabrication shop. Through close relationships with our customers-who include some of the most innovative and largest MEP contractors-we have developed a suite of Stratus tools to digitize, automate, and optimize piping, plumbing, sheet metal, and electrical contracting. Stratus provides the software layer an MEP Contractor needs to optimize profits with true “Data Driven Contracting.” GENERAL DESCRIPTION

The Senior Data Architect owns our canonical data architecture - the schema, contracts, tenancy, and governance that every product and every AI/ML workload builds on. You are the single owner of the canonical data model: one normalized definition of the core business objects shared across our products, and the standard the rest of engineering builds against. This is a foundational, hands-on role - you design, prototype, and ship reference implementations and in-repo guardrails, not just diagrams.

Our approach to AI is to build durable, domain-specific data assets rather than commodity model infrastructure: we don’t pretrain foundation models and we don’t ship thin wrappers around someone else’s. The differentiated value lives in how our data is modeled, governed, and made trustworthy for AI - and that is the layer you own. KEY RESPONSIBILITIESAI/ML readiness

  • Architect the data layer so AI/ML workloads - vector search, embeddings pipelines, RAG-grounded retrieval, model training - run on a clean, governed substrate.
  • Make production data AI-ready: well-modeled, contract-enforced, lineage-tracked, and drift-detectable.
  • Design the data-side integration patterns these workloads depend on, such as feature-store and vector-store patterns across document, relational, and embedding data.

Data architecture

  • Own the canonical data model - the normalized definition of the core business objects shared across our products - and decide what is canonical versus tenant-specific.
  • Establish data architecture standards, data contracts, and schema discipline the rest of engineering builds against, enforced in-repo.
  • Exercise strong polyglot-persistence judgment: what belongs in document vs. relational vs. vector stores, and how to migrate between them without big-bang rewrites.
  • Define the multi-tenant data architecture: tenancy isolation, data residency posture, and per-tenant cost attribution across storage and compute.

Modernization

  • Lead staged modernization toward the right mix of stores and patterns for transactional, analytical, and AI/ML use cases - improving scalability, governance, and usability while minimizing disruption.
  • Own the architectural direction of the data pipeline and lake / lakehouse layer: ingestion, transformation, orchestration, and storage tiers.
  • Lead the move from homegrown pipelines to proven, industry-standard platforms, balancing build-vs-buy and total cost of ownership.
  • Modernize legacy data-access patterns via incremental, strangler-fig migrations that keep production stable.

Technical leadership

  • Drive hands-on prototypes, reference implementations, and in-repo guardrails.
  • Define the data, storage, and retrieval patterns the rest of engineering builds against.
  • Establish data quality, testing, lineage, and observability standards for pipelines and AI/ML serving.
  • Mentor engineers on schema discipline, modern data practices, and AI/ML-readiness patterns.
  • Make canonical decisions that are time-boxed, written, and defensible; hold disagree-and-commit rather than letting schema debate become a standing committee.
  • Use AI-assisted development tools (Claude Code, Copilot, Cursor) as a force multiplier for schema design, query tuning, and migration scripting.

Cross-team partnership

  • Partner with database engineering on production data health while owning long-term architectural direction.
  • Partner with ML and application engineering on their data needs - structuring and governing data so it is retrieval-ready and safe to build on.
  • Partner with platform / infrastructure on reliability, disaster recovery, residency, and the multi-tenant operational posture., * The canonical data model is owned and enforced: teams build against stable, documented contracts instead of bespoke forks.
  • Workloads sit in the right stores, legacy anti-patterns are receding, and reliability targets are holding.
  • Tenancy is formalized and per-tenant cost attribution is instrumented, so cost and capacity are observable as we scale.
  • The data substrate is AI-ready - model, contracts, and lineage in place - so AI/ML work builds on a solid foundation rather than waiting on data.
  • You’ve done it in partnership: the data tier is healthier, and engineers build against your contracts., Senior hands-on tech lead/solution architect for Banking & Capital Markets driving client relationships, solution design, proposals, and delivery of data modernization, streaming, and cloud-based data platforms. Leads cross-functional teams, mentors staff, supports revenue growth through upsell/cross-sell, and provides industry thought leadership and technical strategy. Top Skills: Ai/Ml ToolsAPIsAWSAzureBig Data TechnologiesDatabricksElasticsearchETLGCPIn-Memory Caching DatabasesKafkaMongoDBNosql DatabasesSnowflakeSQLStreaming Technologies Inspira Financial, Owns strategy, governance, and execution of knowledge content across Intercom, NICE CXone Expert, and Guru. Sets taxonomy, style, and lifecycle standards; drafts, edits, and publishes articles; manages SME and compliance reviews; runs platform syncs and bot gap reviews; tracks content performance and readiness for product and policy launches; and builds the case to scale the team and transition execution to a dedicated Content Specialist. Top Skills: Ai-Assisted Customer Service ToolsChatbotsConfluenceGuruIntercom Knowledge HubNice Cxone ExpertVirtual AgentsZendesk Inspira Financial

Health & Benefits - AI Bot Optimization Specialist (Remote)

32 Minutes Ago In-Office or Remote 36-40 Hourly Junior 36-40 Hourly Junior Fintech Own monitoring, testing, and tuning of Inspira’s AI virtual agent to improve resolution, deflection, and CSAT. Analyze conversation transcripts and performance data, run A/B and regression tests, adjust retrieval/configuration and intents, partner with Content and IT to fix knowledge gaps, maintain performance dashboards, lead root-cause analysis, and support go-lives for new bot capabilities. Top Skills: A/B TestingAi Virtual Agent PlatformsCx Analytics (CsatFin VoiceIntercomKnowledge Base SystemsNpsPrompt EngineeringResolution/Deflection Metrics)Retrieval-Augmented Generation (Rag)

What you need to know about the Colorado Tech Scene

With a business-friendly climate and research universities like CU Boulder and Colorado State, Colorado has made a name for itself as a startup ecosystem. The state boasts a skilled workforce and high quality of life thanks to its affordable housing, vibrant cultural scene and unparalleled opportunities for outdoor recreation. Colorado is also home to the National Renewable Energy Laboratory, helping cement its status as a hub for renewable energy innovation.

Key Facts About Colorado Tech

  • Number of Tech Workers: 260,000; 8.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Lockheed Martin, Century Link, Comcast, BAE Systems, Level 3
  • Key Industries: Software, artificial intelligence, aerospace, e-commerce, fintech, healthtech
  • Funding Landscape: $4.9 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Access Venture Partners, Ridgeline Ventures, Techstars, Blackhorn Ventures
  • Research Centers and Universities: Colorado School of Mines, University of Colorado Boulder, University of Denver, Colorado State University, Mesa Laboratory, Space Science Institute, National Center for Atmospheric Research, National Renewable Energy Laboratory, Gottlieb Institute

Requirements

  • 8+ years in data architecture, data engineering, database administration, or analytics engineering, with 3+ years in senior / lead roles.
  • Demonstrated ownership of a canonical or enterprise data model / cross-product schema - the model and contracts other teams built against.
  • Hands-on MongoDB at production scale (Atlas M40+ ideal): document modeling, aggregation framework, indexing, change streams, sharding, replica sets - and the judgment to recognize the Mongo-as-RDBMS anti-pattern.
  • Strong polyglot-persistence judgment: deciding what belongs in documents vs. relational vs. a vector store, and migrating between them incrementally.
  • Hands-on relational depth: schema design, indexing strategy, and query tuning, plus familiarity with vector search (Atlas Vector Search, pgvector, or equivalent).
  • Production experience making data AI/ML-ready: data architecture supporting RAG, semantic search, embeddings / vector pipelines, or agentic workloads.
  • Multi-tenant architecture experience: data residency and per-tenant cost attribution.
  • Pipeline / ELT / lake / lakehouse design at scale, with incremental migration strategies that minimize disruption.
  • Cloud-native data services (Azure, AWS, or GCP).
  • Strong grasp of data quality, testing, lineage, and monitoring - including observability for pipelines and AI/ML serving.
  • Comfortable modeling a complex, specialized domain. MEP / AEC / construction experience is a plus; appetite to learn the domain is required.

NICE TO HAVE

  • Knowledge-graph, ontology, or semantic-layer experience.
  • CDC and cross-engine sync (MongoDB Change Streams, Debezium, or equivalent).
  • Lakehouse platforms (Databricks, Snowflake, or open table formats - Iceberg, Delta, Hudi) and feature stores (Feast or equivalent).
  • Data governance for AI/agent access to production data: query-cost controls, read-path safety, lineage, and audit for higher-risk use cases.
  • SOC 2 and data-classification experience.
  • Azure data ecosystem (Data Factory, Synapse, Functions, Event Grid).
  • MongoDB certification (Associate DBA / Developer or higher) or substantive MongoDB University coursework.

Benefits & conditions

  • Comprehensive and competitive health benefits plan
  • Matching 401k contributions
  • 20 days annual PTO
  • Primarily remote work with occasional annual team onsites

This is a fully remote position open to candidates based in the United States.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on ats.rippling.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

1:26 min

Building an initial solution using Debezium and Apache Kafka

Bobur Umurzokov · LIVE

2:01 min

Migrating existing applications from MongoDB to Postgres

Nikita Shamgunov Nikita Shamgunov · WWC 2024

2:36 min

Analyzing limitations with PostgreSQL bitmap heap scans

Dharin Shah Dharin Shah · WWC 2025

2:59 min

Solving data ingestion and recognizing tool boundaries

Matthias Niehoff Matthias Niehoff · WWC 2023

1:42 min

Handling stateful application schemas during progressive deployment

Kevin Dubois Kevin Dubois · WWC 2025

Videos

See all

Related articles

See all