> Markdown version of [/jobs/ext/2462950-principal-software-developer-data](https://www.wearedevelopers.com/jobs/ext/2462950-principal-software-developer-data). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Software Developer - Data - **Company:** Rocket Companies - **Location:** Seattle, WA, United States - **Experience:** Expert - **Salary:** $234,000.0 - $286,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Amazon S3, Apache HTTP Server, Business Logic, Microsoft Azure, Cloud Computing, Databases, Continuous Integration, Data Architecture, Information Engineering, Data Governance, Data Infrastructure, Data Integration, Extract Transform Load (ETL), Data Sharing, Data Warehousing, Dimensional Modeling, Event Logging, Github, Identity and Access Management, Python (Programming Language), Machine Learning, Open Source Technology, Role-Based Access Control, SQL Databases, Data Streaming, Enterprise Data Management, Data Ingestion, Large Language Models, Snowflake, Apache Spark, Pyspark, Infrastructure Automation Frameworks, Star Schema, Apache Kafka, Machine Learning Operations, Data Lakehouse, Terraform, Automation Anywhere, Web Api - **Published:** August 5, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=e47a95ac478c93e6 ## About the Role * Experience: 12+ years of software/data engineering experience, with at least 5+ years operating at a Staff, Principal, or Lead Architect level. * Cloud & Lakehouse Expertise: Deep, hands-on architectural experience with AWS (S3, IAM, EKS) and Snowflake. Proven experience implementing open table formats, specifically Apache Iceberg. * Data Ingestion & ELT: Deep expertise with modern data movement pipelines (e.g., Fivetran, Airbyte) and high-velocity event-streaming platforms (e.g., Apache Kafka, AWS Kinesis, Snowpipe). Experience migrating legacy ETL platforms to modern, compute-optimized ELT architectures is highly preferred. * Data Modeling: Expert understanding of modern data modeling techniques, including Star Schema design for highly optimized analytical reads and dimensional modeling for enterprise semantic layers. * Processing Frameworks: Advanced proficiency in SQL, Python, and distributed compute engines (Spark, Ray, or native Snowflake compute). * Data Governance & Security: Extensive experience designing multi-tenant data architectures, implementing complex Role-Based Access Control (RBAC), and managing cross-domain security policies in highly regulated environments. * CI/CD & IaC: Mastery of CI/CD pipelines (e.g., Azure DevOps, GitHub Actions) and infrastructure provisioning tools (Terraform). The Ideal Candidate Will Also Have: * A strong perspective on the evolving "Modern Data Stack" and the tradeoffs between tight vendor coupling versus open-source flexibility. * Experience integrating data platforms directly with Machine Learning pipelines and AI workflows. * A proven track record of influencing engineering culture and driving alignment across multiple disparate engineering squads. ## Description We are seeking a Principal Software Engineer to serve as the chief architect and technical north star for our enterprise data organization. In this role, you will design and scale a unified data platform that serves as the central nervous system for the entire company. You will lead the transition to a modern, multi-tenant Data Lakehouse, abstracting away infrastructure complexity for our data engineers while democratizing high-performance, self-serve access for data scientists, machine learning engineers, and AI agents. You will partner with engineering directors and mentor Staff-level ICs to ensure our architecture scales elegantly securely across domain boundaries., * Lakehouse Architecture & Strategy: Design and implement a unified, open-format Data Lakehouse utilizing a Snowflake-managed Apache Iceberg architecture on AWS S3, transitioning the organization away from disjointed legacy data silos. * Enterprise Data Ingestion: Architect and scale resilient, high-throughput ingestion frameworks. Design batch, micro-batch, and real-time streaming patterns to seamlessly extract data from diverse transactional databases, third-party APIs, and event logs, landing it reliably into our Bronze storage tier. * Multi-Tenant Data Mesh Design: Architect a sovereign, domain-oriented data mesh within a single Snowflake environment. Implement delegated domain administration and sophisticated RBAC/Row-Access Policies to ensure strict infosec compliance without impeding cross-domain data sharing. * Semantic Layer & AI Readiness: Spearhead the deployment of a Universal Semantic Layer on top of our Gold-tier data. Ensure business logic is defined as code to guarantee metric consistency across BI tools, data apps, and downstream LLM/AI agents. * Data Flow & Pipeline Engineering: Define the technical standards for our Medallion (Bronze, Silver, Gold) data flow. Standardize modern ELT patterns utilizing dbt, PySpark, and Snowpark to handle petabyte-scale transformations efficiently. * Infrastructure & Automation: Drive a rigorous Infrastructure-as-Code (IaC) culture using Terraform for all platform provisioning, networking, and security configurations. * Technical Leadership: Act as the ultimate technical escalation point for an organization of 60+ engineers. Mentor Staff and Senior Individual Contributors, lead architecture design reviews, and establish engineering best practices. * Cross-Organization Collaboration: Act as the primary technical liaison to Principal Architects across the broader organization. Define robust integration contracts and connection points between upstream transactional systems, the data platform, and downstream product applications, ensuring a unified, highly interoperable enterprise architecture., This role may include participation in an on-call rotation to support production systems and ensure service reliability. On-call responsibilities may include coverage during nights and weekends. If applicable, frequency and scheduling will be determined by team needs and communicated accordingly. ## Related Videos - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Web APIs you might not know about](https://www.wearedevelopers.com/videos/281-web-apis-you-might-not-know-about) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [What is Software Engineering in the Age of AI?](https://www.wearedevelopers.com/magazine/640-what-is-software-engineering-in-the-age-of-ai)