Principal Software Engineer - Data Platform

Innovaccer, Inc.
United States
2 days ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Amazon S3 Apache HTTP Server Big Data BigQuery Cloud Database Distributed Systems Apache Hive Python (Programming Language) Open Source Technology Software Engineering SQL Databases
+8 more
Snowflake Apache Spark Change Data Capture Data Lakes Kubernetes Information Technology Data Management Presto

Job description

As a Principal Software Engineer on the Data Platform, you will own the architecture of the lakehouse engine that powers Innovaccer’s platform in on-premise deployments: Apache Iceberg as the table format, Trino as the query engine, a REST catalog service, and Spark as transform compute. Cloud data warehouses have no on-premise equivalent, so this is a ground-up engine design, not a re-point. It is the single longest-lead technical track in the program, and the decisions you make on catalog, engine placement, and pipeline redesign gate everything downstream: transforms, serving, and reporting.

A Day in the Life

  • Own the lakehouse reference architecture: Iceberg table design, Trino cluster topology, catalog service, Spark transform compute, and object-storage layout.
  • Design on-premise replacements for cloud-managed warehouse capabilities that have no direct equivalent: change-data-capture streams, scheduled tasks, and write-back paths into operational stores.
  • Run proof-of-concept validation of the catalog and query engine at expected data volumes, and define evidence-based triggers for placement decisions (VM-based versus Kubernetes-native operators).
  • Set platform-wide standards for table layout, partitioning, file sizing, and Iceberg maintenance: compaction, snapshot expiry, and orphan-file cleanup.
  • Lead the SQL dialect strategy for porting existing warehouse workloads to Trino and Spark SQL.
  • Mentor senior engineers across data workstreams, review designs, and raise the bar on engineering quality.
  • Partner with platform engineering on storage sizing, resource isolation, and capacity planning for the lakehouse footprint.

Requirements

  • B.E., B.Tech., M.Sc. degree in Computer Science or a related technical field.
  • 12+ years of industry experience building and operating large-scale data platforms or distributed systems.
  • Deep, hands-on expertise with distributed SQL engines: Trino/Presto or Spark SQL internals, query planning, and performance engineering.
  • Production experience with Apache Iceberg (or Delta Lake/Hudi with willingness to go deep on Iceberg): table spec, merge-on-read versus copy-on-write, and table maintenance at scale.
  • Working knowledge of Iceberg catalog services (REST catalogs such as Polaris or Nessie, or Hive Metastore) and S3 compatible object storage.
  • Strong understanding of cloud warehouse internals (Snowflake, BigQuery, or Redshift) sufficient to design functional equivalents on open-source infrastructure.
  • Professional software development experience with Java and/or Python.
  • Experience delivering data platforms in on-premise, regulated, or air-gapped
  • environments is a strong plus; healthcare data experience is a plus.

Benefits & conditions

We offer competitive benefits to set you up for success in and outside of work., * Generous Paid Time Off: Recharge and relax with 20 days of fixed time off per year, in addition to company holidays-because we believe work-life balance fuels performance.

  • Best-in-Class Parental Leave: Spend quality time with your growing family. We offer one of the industry’s most generous parental leave policies to support you during life’s most important moments.
  • Recognition & Rewards: We celebrate wins-big and small. Get rewarded with monetary incentives and company-wide recognition for your impact and dedication. Your hard work won’t go unnoticed.
  • Comprehensive Insurance Coverage: Stay covered with medical, dental, and vision insurance, plus 100% company-paid short- and long-term disability and basic life insurance. Optional perks include discounted legal aid and pet insurance.

Innovaccer Inc. is an equal opportunity employer. We celebrate diversity and are committed to fostering an inclusive workplace where all employees feel valued and empowered regardless of any characteristic protected by federal, state or local law including, without limitation, race, color, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, medical condition, disability, age, marital status, or veteran status. Innovaccer Inc. participates in the E-Verify program to confirm employment eligibility of all newly hired employees based out of the U.S. and employed by Innovaccer Inc.

About the company

Innovaccer activates the flow of healthcare data, empowering providers, payers, and government organizations to deliver intelligent and connected experiences that advance health outcomes. The Healthcare Intelligence Cloud equips every stakeholder in the patient journey to turn fragmented data into proactive, coordinated actions that elevate the quality of care and drive operational performance. Leading healthcare organizations like CommonSpirit Health, Atlantic Health, and Banner Health trust Innovaccer to integrate a system of intelligence into their existing infrastructure, extending the human touch in healthcare. For more information, visit www.innovaccer.com.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou ¡ Coffee With Developers

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy ¡ LIVE

2:30 min

Leveraging BigQuery ML for scalable SQL-based segmentation experiments

Julian Joseph ¡ LIVE

2:14 min

Solving complex platform architecture challenges at an enterprise scale

Maria Apazoglou ¡ Coffee With Developers

2:10 min

Why organizations combine big data and machine learning

Ayon Roy ¡ LIVE

3:27 min

Explaining query execution overhead and caching limitations in BigQuery

Adnan Rahic ¡ JS Congress

Videos

See all

Related articles

See all