Staff Software Engineer, Data Warehouse
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+18 more
Job description
We’re hiring a Staff Software Engineer to own Commure’s data warehouse platform end-to-end. You’ll design, build, and operate every layer of the stack:
- CDC pipelines
- Streaming transport
- Schema governance and data contracts
- Query and serving layer
- Analytics platform
The core stack today: Debezium for CDC, StarRocks as our MPP query and serving engine, and dbt for transformation and modeling. This is a hands-on IC role with broad scope. You’ll make architectural calls, write the code that matters most, and set the patterns other teams build on.
What You’ll Do
- Own the data warehouse platform end-to-end: CDC pipelines, data lake, query layer, transformation layer, and the analytics-facing tooling that sits on top.
- Design and operate CDC pipelines with Debezium (and Kafka, Redpanda, or an equivalent streaming backbone) that move data from operational databases into the warehouse with low latency and high fidelity.
- Architect the data lake on object storage using an open table format (Iceberg, Delta Lake, or Hudi) with Parquet, enabling both batch and streaming workloads and clean separation of storage from compute.
- Run and scale StarRocks (or adjacent MPP/lakehouse engines) as the query and serving layer - schema design, materialized views, ingestion patterns, tuning, and cost/performance trade-offs.
- Build the transformation layer with dbt: modeling standards, tests, documentation, and a semantic layer that gives every team a single source of truth for metrics.
- Stand up orchestration (Airflow, Dagster, or similar) and the CI/CD, observability, and data-quality tooling that make the platform trustworthy day-to-day.
- Partner with Security and Compliance on PHI/PII handling, access controls, lineage, and auditability so the platform meets HIPAA and SOC 2 bar by default.
- Set patterns and conventions: schema contracts, ingestion patterns, and self-serve tooling - that let product and analytics teams build on the platform without needing you in the loop for every decision.
Requirements
- 6+ years of software engineering experience, with significant time building or operating data platforms at scale.
- Experience across the modern data stack: CDC (Debezium or equivalent), streaming (Kafka, Redpanda), data lake formats (Iceberg, Delta, Hudi), an MPP or lakehouse query engine (StarRocks, ClickHouse, Trino, Snowflake, Databricks), and dbt.
- Fluent in SQL, schema design, query optimization, and reasoning about cost and latency trade-offs on large datasets.
- Experience running production data infrastructure (orchestration, observability, on-call, data quality, and incident response).
Preferred
- Direct experience with Debezium, StarRocks, and dbt in production.
- Experience building semantic layers (dbt Semantic Layer, Cube) or data catalogs / lineage (DataHub, OpenMetadata, Amundsen).
- Experience with HIPAA-regulated data (PHI handling, de-identification, and access governance).
- Experience powering AI/ML workloads: feature stores, training-set curation, embedding pipelines, or retrieval systems.
- Experience across multiple clouds (AWS, GCP, Azure), infrastructure-as-code (Terraform, Pulumi) and Kubernetes controllers.
About the company
At Commure, we’re building the AI Operating System for healthcare, the foundation that defines how care is delivered, documented, and financed. Our platform spans the full care journey: Ambient AI and Dictation eliminating documentation burden at the point of care, intelligent Agents automating patient and revenue workflows, and autonomous RCM processing billions in claims, all on a single AI-native platform integrated with 60+ EHRs.
Healthcare carries a $1 trillion administrative burden and we’re at the center of transforming it. Today, 500,000+ clinicians across 500+ healthcare organizations nationwide trust Commure to handle $25B+ in annual claims and support over 200 million patient interactions. Our latest $70M raise at a $7B valuation reflects the confidence the market has placed in this mission. We’ve also been named to the Fortune Future 50 list and the 2026 AI Breakthrough Awards for “Overall NLP Company of the Year.”
Our team works directly alongside clinicians, not through layers of process, which means the gap between what you build and its impact on patient care is immediate. We move fast, deploy daily, and take full ownership from early thinking to production. If you’re energized by hard problems, high stakes, and a team that holds itself to a high bar, you’ll find your people here.
The future of healthcare is being built right now. Come deliver this transformation.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
Highest Paying Tech Companies for Developers
Dev Digest 120 - Apple and peers
Top Big Data Technologies That You Need to Know