> Markdown version of [/jobs/ext/2715059-principal-software-engineer-data-platform](https://www.wearedevelopers.com/jobs/ext/2715059-principal-software-engineer-data-platform). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Software Engineer - Data Platform - **Company:** Innovaccer, Inc. - **Location:** United States (Remote available) - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Amazon S3, Apache HTTP Server, Big Data, BigQuery, Cloud Database, Distributed Systems, Apache Hive, Python (Programming Language), Open Source Technology, Software Engineering, SQL Databases, Snowflake, Apache Spark, Change Data Capture, Data Lakes, Kubernetes, Information Technology, Data Management, Presto - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/principal-software-engineer-data-platform-iceberg-trino-innovaccer-analytics-9000758 ## About the Role * B.E., B.Tech., M.Sc. degree in Computer Science or a related technical field. * 12+ years of industry experience building and operating large-scale data platforms or distributed systems. * Deep, hands-on expertise with distributed SQL engines: Trino/Presto or Spark SQL internals, query planning, and performance engineering. * Production experience with Apache Iceberg (or Delta Lake/Hudi with willingness to go deep on Iceberg): table spec, merge-on-read versus copy-on-write, and table maintenance at scale. * Working knowledge of Iceberg catalog services (REST catalogs such as Polaris or Nessie, or Hive Metastore) and S3 compatible object storage. * Strong understanding of cloud warehouse internals (Snowflake, BigQuery, or Redshift) sufficient to design functional equivalents on open-source infrastructure. * Professional software development experience with Java and/or Python. * Experience delivering data platforms in on-premise, regulated, or air-gapped * environments is a strong plus; healthcare data experience is a plus. ## Description As a Principal Software Engineer on the Data Platform, you will own the architecture of the lakehouse engine that powers Innovaccer's platform in on-premise deployments: Apache Iceberg as the table format, Trino as the query engine, a REST catalog service, and Spark as transform compute. Cloud data warehouses have no on-premise equivalent, so this is a ground-up engine design, not a re-point. It is the single longest-lead technical track in the program, and the decisions you make on catalog, engine placement, and pipeline redesign gate everything downstream: transforms, serving, and reporting. A Day in the Life * Own the lakehouse reference architecture: Iceberg table design, Trino cluster topology, catalog service, Spark transform compute, and object-storage layout. * Design on-premise replacements for cloud-managed warehouse capabilities that have no direct equivalent: change-data-capture streams, scheduled tasks, and write-back paths into operational stores. * Run proof-of-concept validation of the catalog and query engine at expected data volumes, and define evidence-based triggers for placement decisions (VM-based versus Kubernetes-native operators). * Set platform-wide standards for table layout, partitioning, file sizing, and Iceberg maintenance: compaction, snapshot expiry, and orphan-file cleanup. * Lead the SQL dialect strategy for porting existing warehouse workloads to Trino and Spark SQL. * Mentor senior engineers across data workstreams, review designs, and raise the bar on engineering quality. * Partner with platform engineering on storage sizing, resource isolation, and capacity planning for the lakehouse footprint. ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Making Data Warehouses fast. A developer's story.](https://www.wearedevelopers.com/videos/302-making-data-warehouses-fast-a-developer-s-story) - [Data Governance in the Era of AI](https://www.wearedevelopers.com/videos/1622-data-governance-in-the-era-of-ai) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)