Principal Software Engineer, Real Time Data Enrichment Platform (Hybrid)

CrowdStrike
New York, NY, United States
15 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours

Tech stack

Query Performance Java (Programming Language) Artificial Intelligence Apache HTTP Server Big Data Data Infrastructure Distributed Systems Memory Management Amazon DynamoDB Apache Hadoop Apache Hive Java Virtual Machine (JVM)
+28 more
PostgreSQL Linux System Administration Machine Learning MySQL Online Analytical Processing NoSQL Performance Tuning Standard Sql Mesos Scala (Programming Language) Software Engineering Data Streaming Data Storage Technologies Real Time Systems Test-Driven Development (TDD) Apache Spark Indexer Kotlin Data Lakes Kubernetes Druid Apache Flink Cassandra Data Analytics Real Time Data Apache Kafka Data Management Presto

Job description

As a global leader in cybersecurity, CrowdStrike protects the people, processes and technologies that drive modern organizations. Since 2011, our mission hasn’t changed - we’re here to stop breaches, and we’ve redefined modern security with the world’s most advanced AI-native platform. We work on large scale distributed systems, processing almost 3 trillion events per day and this traffic is growing daily. Our customers span all industries, and they count on CrowdStrike to keep their businesses running, their communities safe and their lives moving forward. We’re proud to work for a mission-driven company leveraging AI to transform the way we work. CrowdStrikers drive their careers through flexibility and autonomy while also being expected to contribute to a culture of responsible AI adoption, experimentation, and innovation. We use an AI-first mindset as a force multiplier to proactively and continuously accelerate execution, build expertise, uncover insights, and solve complex problems. We’re always looking to add talented CrowdStrikers to the team who have limitless passion, a relentless focus on innovation and a fanatical commitment to our customers, our community and each other. Ready to join a mission that matters? The future of cybersecurity starts with you.

About the Role:

Our Data Platform group at Crowdstrike is unique among its kind for being uncommonly customer-focused. We build and operate systems to centralize all of the data from Falcon Sensors and 3rd Party sources derived from trillions of events/day, and we also drive industry-leading innovation on a hyper scale security data lake that helps find bad actors and stop breaches.

Setting ourselves apart, we make it easy for all customers to utilize the platform for batch and streaming analytics, machine learning and threat hunting through the production and delivery of self-service platforms (Query Platform, Analytics Platform, Enrichment platform…) using Spark and Flink respectively as the runtimes. These self-service platforms allow customers to apply their own custom schema, syntax, data models (etc.) to our historical cyberattack data, modeling threats in a way that empowers them to predictively build (among other things) behavioral automations that defend against such threats ‘before’ they appear in their own environments.

In this role, you will be a Principal Engineer within Data Platform, owning - in entirely hands-on capacity the design, build and delivery of a new self-service data enrichment platform which will take our customers’ proactive defense measures to the next level.

As a leader in data platform, you will contribute to the full spectrum of our systems, including query processing, scalable pipeline builds with largely Apache-based ingestion, materialized view, transformation and data storage frameworks, and tools/applications that make data available to thousands of users and hundreds of internal systems.

What You’ll Do:

  • Shape the vision of our Analytics Data Platform for its next phase of growth: building a Unified Data Catalog as well as Query Analysis for structured data stored in different forms (columnar vs graph), and building and optimizing query performance using different techniques of indexing and data partitioning.
  • Design, develop, and maintain a data platform that processes petabytes of data.
  • Participate in technical reviews of our products and help us develop new features and enhance stability.
  • Continually help us improve the efficiency of our services so that we can delight our customers.
  • Help us research, evolve and implement new ways for both internal stakeholders as well as customers to query their data efficiently and extract results in the format they desire.

Requirements

  • (One among:) 17+ years exp with B.S. in a related field, 15+ years with M.S. in a related field, or 12+ years with PhD in a related field.
  • Experience building and supporting very high scale data platform and data storage systems (MINIMUM: 100s of TB/day in either a current or past role)
  • Significant experience performance-tuning or developing internals (source code) within Spark, Flink, Iceberg or Pinot or an equivalent structured streaming, Time Series, OLAP or Open Table real time system.
  • Production experience building either Spark- or Flink-based self-service data platforms, or equivalent (i.e. with Ray, or building spark- or Flink-like frameworks themselves in Scala, Akk, etc.)
  • Strong familiarity with (and ample hands-on experience tuning & optimizing) at least one applicable technology in the Apache Hadoop ecosystem: Spark, Kafka, Hive/Iceberg/Delta Lake, Presto/Trino, Pinot, Druid, etc.
  • 3+ years coding in Java, Scala, Kotlin or another JVM language (bonus points for experience tuning the language, i.e. garbage collection, memory management…)
  • Production experience with relational SQL and NoSQL databases, including Postgres/MySQL, Cassandra/DynamoDB, etc.
  • Proven expertise with multiple big data frameworks in general, especially handling data volume at (ideally) multi-petabyte scale.
  • Proven expertise with algorithms, distributed systems design and the software development lifecycle.
  • Great test driven development discipline.
  • Reasonable proficiency with Linux administration tools.
  • Proven experience utilizing AI technologies to enhance decision-making, streamline workflows and processes, improve efficiency and drive business outcomes.
  • Proven ability to work effectively with remote teams.

Bonus Points:

  • Familiarity with Go.
  • Familiarity with Kubernetes/Mesos or equivalent.
  • Production experience with Flink, especially as runtime for a self-service platform.

LI-MP2

About the company

Benefits of Working at CrowdStrike:

  • Market leader in compensation and equity awards
  • Comprehensive physical and mental wellness programs
  • Competitive vacation and holidays for recharge
  • Paid parental and adoption leaves
  • Professional development opportunities for all employees regardless of level or role
  • Employee Networks, geographic neighborhood groups, and volunteer opportunities to build connections
  • Vibrant office culture with world class amenities
  • Great Place to Work Certified across the globe

CrowdStrike is proud to be an equal opportunity employer. We are committed to fostering a culture of belonging where everyone is valued for who they are and empowered to succeed. We support veterans and individuals with disabilities through our affirmative action program.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on crowdstrike.wd5.myworkdayjobs.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:29 min

Evaluating orchestration tools for distributed container deployments

Aleksandr Kalikov · LIVE

2:37 min

Comparing traditional SQL tables versus NoSQL non-tabular databases

Stanimira Vlaeva · JS Congress

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

2:14 min

Solving complex platform architecture challenges at an enterprise scale

Maria Apazoglou · Coffee With Developers

3:16 min

Terminology differences between relational and NoSQL databases

Tim Faulkes · LIVE

5:53 min

Choosing relational databases over NoSQL for most workloads

Josip Stuhli Josip Stuhli · WWC Europe 2026

Videos

See all

Related articles

See all