Skip to content

Real-Time Data Platforms at Trillion-Event Scale

with Diptamay Sanyal

Thursday 24 September 2:10 PM – 2:40 PM Stage 3

About This Session

Some of the largest real-time platforms in production today operate at scales that were considered theoretical a decade ago. CrowdStrike has publicly disclosed that its Falcon Threat Graph processes more than a trillion events per day across 15-plus petabytes of data, with engineering blogs describing 40-plus petabytes stored and 70 million requests per second served. Numbers like these are not unique to security — AI agent platforms, observability backends, and large SaaS analytics systems are pushing into the same territory. The interesting question is not the headline figure. It is what actually breaks at that scale, and which architectural choices keep the system honest. This talk distills patterns and failure modes from years of building streaming and AI data platform infrastructure, framed against publicly disclosed industry references rather than any single employer's internals. Topics include: - Stateful streaming joins across different data streams - Handling replay storms and out-of-order events without latency collapse - The honest tradeoffs between latency, cost, and correctness — and when each one wins - Observability and degradation modes that keep the platform usable when something is failing Attendees will leave with concrete guidance on designing real-time systems that fail loudly, recover predictably, and surface the metrics that actually matter at scale. Note: Views are my own and do not represent any employer. Examples reference publicly disclosed sources.

Topics

  • Apache Flink
  • Apache Kafka
  • Agentic AI
  • Data Pipelines
  • Large Language Models (LLMs)