Staff Data Architect

JELLYFISH, LLC
United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$200,000.0 - $260,000.0
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Airflow BigQuery Cloud Storage Databases Data Architecture Data Validation Data Governance Data Infrastructure DevOps Middleware Python (Programming Language)
+21 more
Automation of Marketing Operational Databases Raw Data DataOps Software Engineering Management of Software Versions Sql Optimization ReactJS Large Language Models Snowflake Technical Debt Backend Git Core Data Data Lineage Data Management Front End Software Development Terraform Code Restructuring Data Pipelines Databricks

Job description

Jellyfish is the backbone for elite engineering organizations, and our data infrastructure needs to be as high-performing and insightful as the teams we serve. We are looking for a Staff/Lead Data Architect to help us design, automate, and scale the next generation of our Jellyfish data platform. You’ll be responsible for maturing our core data models, automating environment boundaries, and driving advanced observability and cost-attribution deeper into our data pipeline architecture. If you view manual data intervention as a technical debt to be solved and want to work in an environment where your architectural decisions directly impact how the world’s best engineering leaders measure their productivity, you’re the perfect fit., * Architectural Evolution & Blueprinting - You’ll own the blueprint for the next-generation Jellyfish data platform. You’ll tackle our existing data footprint, refactoring pipelines and structures into highly efficient, scalable patterns (like Medallion-style schemas or unified semantic layers).

  • Automated Data Governance - You’ll design and automate strict, code-driven environment isolation boundaries. You’ll ensure dev, staging, and production data catalogs (and their underlying cloud storage) never dangerously cohabitate, eliminating the risk of “fat-finger” data drops or PII leakage.
  • Orchestration & Compute Scaling - You’ll lead the modernization of our workflow orchestration and distributed compute engines. You’ll focus on slashing engine runtime overhead, eliminating API bottlenecks, and streamlining heavy parallelized or mapped data tasks.
  • Modern Integration Middleware - You’ll partner with application teams to ensure our React frontends and backend services hit highly secure, cached API and Backend-for-Frontend (BFF) layers rather than querying raw data services directly, protecting our warehouses from concurrency spikes.
  • Proactive Data Observability & FinOps - You’ll build and maintain granular data-quality monitors and cost-allocation frameworks. You won’t just track overall warehouse spend; you’ll implement systems to map execution cost and token usage directly down to the tenant, team, or user level., * You have strong opinions on the future of Git-like data versioning and zero-copy cloning (e.g., Iceberg, Nessie).

Requirements

Do you have experience in Tooling?, * Data Tooling Fluency - You have deep, production-level experience with Python, advanced SQL, and modern data stack essentials. You are deeply familiar with programmatic orchestrators (like Prefect, Dagster, or Airflow) and modern data validation engines (like Pydantic v2).

  • Catalog & Warehouse Practitioner - You have hands-on mastery of enterprise-scale data platforms and governance layers (e.g., Snowflake, Databricks Unity Catalog, BigQuery) and know exactly how to map environments to catalogs and data quality to schemas.
  • Automation Mindset - You look at a manual data backfill or a clicked-together database permission and immediately think about how to automate it via Infrastructure-as-Code (Terraform) or programmatic workflows.
  • Collaborative Systems Thinker - You don’t design in a vacuum. You are excellent at documenting data lineage, mentoring data engineers, and collaborating across DevOps and Product teams to align infrastructure with business goals.
  • Pragmatic Problem Solver - You know the difference between data quality stages and software development lifecycles. You know when a “perfect” distributed cluster is required and when a “good enough” cached view keeps the business moving., * You’ve managed complex cloud-billing attributions or scaled heavy LLM/vector-embedding data workloads and lived to tell the tale.

A list of job experiences and qualification requirements is great, but humility, a performance-driven attitude, and a team-player approach are most important to us. We love to have fun and win in the process. We only hire people who have a passion for building great companies in an environment where a sense of humor is a must.

Occasional travel may be required.

Applicants must be authorized to work for any employer in the US. We are unable to sponsor or take over sponsorship of an employment visa at this time.

Benefits & conditions

3.53.5 out of 5 stars United States Remote $200,000 - $260,000 a year - Full-time

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon ¡ WWC Europe 2026

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer ¡ Coffee With Developers

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum ¡ WWC Europe 2026

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt ¡ LIVE

3:05 min

Audience questions on AI agents and pipeline vectorization

Joy Joy ¡ WWC 2024

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark ¡ LIVE

Videos

See all

Related articles

See all