> Markdown version of [/jobs/ext/2861625-principal-software-architect-data-platform](https://www.wearedevelopers.com/jobs/ext/2861625-principal-software-architect-data-platform). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Software Architect - Data Platform - **Company:** SecurityScorecard - **Location:** United States - **Experience:** Expert - **Salary:** $270,000.0 - $330,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Artificial Intelligence, Amazon Web Services, Batch Processing, Big Data, BigQuery, C++ (Programming Language), Data Architecture, Information Engineering, Data Infrastructure, Data Security, Data Systems, Software Design Documents, Distributed Systems, Python (Programming Language), PostgreSQL, Online Analytical Processing, Node.Js, Data Streaming, Software Technical Review, TypeScript, Parquet, ReactJS, Large Language Models, Snowflake, Apache Spark, Data Layers, Pyspark, Kubernetes, Druid, Apache Flink, Apache Kafka, Machine Learning Operations, Front End Software Development, Vertica, Terraform, Stream Processing - **Published:** September 12, 2026 - **Apply:** https://job-boards.greenhouse.io/securityscorecard/jobs/8193353 ## About the Role * 10+ years of software or data engineering experience, including significant time architecting large-scale data platforms * Deep expertise in stream and batch processing at scale with Kafka, Flink, and Spark or close equivalents, and clear judgment about which path a given workload belongs in * Strong Python and PySpark, solid Java for Flink stream processing, and enough Scala to read and reason about an existing Spark codebase * Hands-on experience designing lakehouse storage in production: columnar formats such as Parquet, open table formats such as Iceberg, and the partitioning, compaction, and schema evolution decisions that come with them * Experience architecting OLAP and analytical serving layers (ClickHouse, Druid, Pinot, BigQuery, Snowflake or similar) * Strong distributed systems fundamentals as they apply to data: exactly-once versus at-least-once semantics, ordering, backpressure, late and out-of-order data, and pipeline failure modes * A track record of building data quality, contracts, and observability as engineered system properties, meaning assertions, schema enforcement, and lineage that live in code * Experience owning a large-scale data migration, including preserving history and correctness through a cutover * A track record of influence without authority: presenting a technical direction to skeptical engineers and earning genuine buy-in, and giving rigorous design review feedback on systems you didn't build yourself * Strong technical writing and mentorship: TDRs, design docs, and decision records that teams can act on without hand-holding, plus a history of raising the technical bar around you * Comfort operating as a senior individual contributor, driving outcomes through prototyping and technical credibility, * Experience building the data layer underneath ML or LLM systems: feature stores, vector stores, or retrieval pipelines * Experience with internet-scale scan, telemetry, or observability data, or a cybersecurity industry background * Track record of bringing platform cost down at scale through storage tiering, query governance, or compute right-sizing, Our platform runs on Node.js and TypeScript, with a React Microfrontend Architecture and PostgreSQL and ClickHouse for storage. We use Kafka for event streaming, and our infrastructure runs on AWS with Kubernetes, Terraform, Helm, and ArgoCD. We are actively expanding our AI/ML infrastructure. On the data side, Kafka is the backbone for event flow. Stream processing runs on Flink in Java, while batch and microbatch run on Spark. Some high performance pipeline components are written in C++. ClickHouse serves our analytical workloads. You do not need to have used every tool here, but you should be comfortable reasoning across a stack of this kind and making principled architectural trade-offs within it. ## Description SecurityScorecard is hiring a Principal Software Architect to lead the system design of our data platform. Rating 12 million companies continuously means ingesting internet-scale measurement data, processing it across streaming, microbatch, and batch paths, storing it so it stays queryable and affordable as it grows, and serving analytics fast enough that customers can explore their own risk in real time. The data is not a byproduct of our product. It is the product. That also raises the stakes on correctness. We publish a number about other companies, they dispute it, and underwriters price against it. A quiet data quality regression here doesn't produce a stale dashboard, it moves someone's score. Quality, contracts, and lineage are therefore architecture problems on this platform, not administrative ones. This is an individual contributor role reporting to the Chief Architect, alongside a Principal Architect focused on AI and agentic systems and a Principal Front-end Architect, and partnering closely with engineering leadership, Product, and Data Science. We're looking for someone who is opinionated about data architecture and persuasive about it: an architect whose designs get adopted because the reasoning is visible, not because they carry a title. Like the rest of our architecture function, you'll prototype to prove out decisions rather than implement full solutions, and you'll set direction through Technical Design Reviews (TDRs), and standards. You will lead the data domain, and you'll bring enough general distributed systems judgment to review designs across the wider platform. What You'll Do * Own the system design for our data platform end to end, from ingestion through to the serving layer * Define the service boundaries and data contracts between producers and consumers, including schema ownership, compatibility rules, and what happens when a producer needs to make a breaking change * Design the lakehouse: table format, partitioning strategy, schema evolution, compaction, and metadata growth at scale * Architect the analytical serving layer for three workload classes with conflicting demands, isolated so that one never degrades another: low-latency, high-concurrency queries from customers in the product, ad-hoc exploration from internal analytics and Data Science, and bulk delivery to external feeds and partners * Set direction on languages and frameworks in the data stack * Engineer data quality and observability into the platform rather than bolting them on: validation and quarantine paths, freshness and completeness SLOs, drift detection, and lineage and metadata generated by the pipeline itself instead of maintained by hand * Design for correctness and reproducibility in the ratings pipeline, including backfills and historical restatement when scoring logic changes * Write the TDRs, design docs, and standards that set data architecture direction across teams, and push that intent into the repos themselves so engineers and coding agents both have it in local context * Review TDRs from across engineering, giving teams substantive feedback on architecture and risk, not only on data work * Partner with the AI & Front End Architects on the data access patterns, mentor senior and staff engineers on data system design, Are you able to work in our NYC office at least two days per week and on an as needed basis?* Select... Will you now or in the future require VISA sponsorship for employment status?* Select... Do you have a Bachelor's degree?* Select... Do you have 10+ years of software/data engineering experience, including significant time architecting large-scale data platforms - regardless of your job title?* Select... Are you strong in Python, with solid Java experience?* Select... Have you worked mainly in an architecture and influence capacity - setting direction and driving adoption across teams - rather than hands-on execution on one team?* Select... Have you designed or guided systems across streaming, microbatch, and batch processing at scale?* Select... Have you spent most of your recent career in enterprise architecture or governance-heavy roles rather than hands-on data platform design?* Select... AI/LLM Usage: Which AI/LLM/agent tools have you used in the last 6 months? For each, note frequency (daily, weekly, occasional) and what you use it for.* AI/LLM Impact: Describe two specific examples where AI improved your work output (speed, quality, clarity, decision-making). Include what you were trying to do, what you asked the tool, and what changed as a result.* Voluntary Self-Identification For government reporting purposes, we ask candidates to respond to the below self-identification survey. Completion of the form is entirely voluntary. Whatever your decision, it will not be considered in the hiring process or thereafter. Any information that you do provide will be recorded and maintained in a confidential file. As set forth in SecurityScorecard's Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law. Gender Select... Are you Hispanic/Latino? Select... Race & Ethnicity Definitions If you believe you belong to any of the categories of protected veterans listed below, please indicate by making the appropriate selection. As a government contractor subject to the Vietnam Era Veterans Readjustment Assistance Act (VEVRAA), we request this information in order to measure the effectiveness of the outreach and positive recruitment efforts we undertake pursuant to VEVRAA. Classification of protected categories is as follows: A "disabled veteran" is one of the following: a veteran of the U.S. military, ground, naval or air service who is entitled to compensation (or who but for the receipt of military retired pay would be entitled to compensation) under laws administered by the Secretary of Veterans Affairs; or a person who was discharged or released from active duty because of a service-connected disability. A "recently separated veteran" means any veteran during the three-year period beginning on the date of such veteran's discharge or release from active duty in the U.S. military, ground, naval, or air service. An "active duty wartime or campaign badge veteran" means a veteran who served on active duty in the U.S. military, ground, naval or air service during a war, or in a campaign or expedition for which a campaign badge has been authorized under the laws administered by the Department of Defense. An "Armed forces service medal veteran" means a veteran who, while serving on active duty in the U.S. military, ground, naval or air service, participated in a United States military operation for which an Armed Forces service medal was awarded pursuant to Executive Order 12985. Veteran Status Select... ## Related Videos - [Parquet, Delta, Iceberg & Ducklake - An introduction for developers](https://www.wearedevelopers.com/videos/100075-parquet-delta-iceberg-ducklake-an-introduction-for-developers) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Tracking vehicles at scale](https://www.wearedevelopers.com/videos/1999-tracking-vehicles-at-scale) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) - [OLAP for AI Applications and why you should care](https://www.wearedevelopers.com/videos/100212-olap-for-ai-applications-and-why-you-should-care) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)