Data Scientist
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
Empirical Security is seeking an experienced Security Data Scientist focused on building the next generation of cybersecurity vulnerability models. Our unique approach leverages ground-truth telemetry to develop predictive, actionable insights that transform the way organizations identify, prioritize, and remediate vulnerabilities in cloud, appsec and traditional environments. We build models specific to individual customers, and maintain many of them side by side. The role
You own models and data end to end: problem framing, features, training, evaluation, deployment, and the uncomfortable part where you explain to a customer why the vulnerability their board is worried about ranked 400th on the remediation list.
What you’ll do
- Design, train, and ship exploit prediction models against ground-truth exploitation telemetry, in cloud, appsec, and traditional infrastructure.
- Build evaluation that survives contact with reality. Precision, recall, coverage, efficiency, calibration, and how all four decay over time. Accuracy is not a number you report once at launch.
- Solve for extreme class imbalance. A fraction of a percent of published CVEs are ever exploited in the wild, and most of the industry’s modeling failures start with pretending that isn’t true.
- Work the hard part of the dual-model architecture: partial pooling, hierarchical priors, and cold-start behavior for customers whose local telemetry is thin in month one and rich in month twelve.
- Engineer features across scanner output, EDR, asset inventory, identity, cloud posture, and exploitation telemetry, and be honest about which ones are leakage.
- Own monitoring and drift detection.
- Publish. Papers, methodology write-ups, open benchmarks, conference talks. Our positioning is that we show our work.
- Partner with engineering and our forward deployed team to move models out of notebooks and into production systems that customers depend on.
Requirements
- Several years of applied machine learning or statistics with models that ran in production and had consequences when they were wrong.
- Fluency in Python and SQL, and the discipline that comes with version control, reproducible pipelines, and secure handling of customer data.
- Real depth in classification under heavy imbalance, plus at least one of: survival and time-to-event analysis, Bayesian hierarchical modeling, or causal inference.
- Calibration instincts. You should be visibly uncomfortable when a model outputs 0.9 and is right 60% of the time.
- The ability to explain a model to a security executive, and to quantify uncertainty out loud rather than burying it in an appendix.
- Enough curiosity about attacker behavior to ask why a feature works, not just whether it does.
A Final Word
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
Résumé-Driven Development: How IT trends affect the job market for software developers
Events like RSAC Get You CISOs. Developers Decide What Actually Gets Deployed.
Trustworthy AI Starts at Deployment: 5 Checks Before You Ship