> Markdown version of [/jobs/ext/2189282-data-scientist](https://www.wearedevelopers.com/jobs/ext/2189282-data-scientist). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Scientist - **Company:** Empirical Security, Inc. - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Customer Data Management, Python (Programming Language), Machine Learning, SQL Databases, Software Version Control - **Published:** August 22, 2026 - **Apply:** https://jobs.ashbyhq.com/empirical-security/f7b15b05-aa9b-4c84-95c5-b5e5917983af ## About the Role * Several years of applied machine learning or statistics with models that ran in production and had consequences when they were wrong. * Fluency in Python and SQL, and the discipline that comes with version control, reproducible pipelines, and secure handling of customer data. * Real depth in classification under heavy imbalance, plus at least one of: survival and time-to-event analysis, Bayesian hierarchical modeling, or causal inference. * Calibration instincts. You should be visibly uncomfortable when a model outputs 0.9 and is right 60% of the time. * The ability to explain a model to a security executive, and to quantify uncertainty out loud rather than burying it in an appendix. * Enough curiosity about attacker behavior to ask why a feature works, not just whether it does. A Final Word ## Description Empirical Security is seeking an experienced Security Data Scientist focused on building the next generation of cybersecurity vulnerability models. Our unique approach leverages ground-truth telemetry to develop predictive, actionable insights that transform the way organizations identify, prioritize, and remediate vulnerabilities in cloud, appsec and traditional environments. We build models specific to individual customers, and maintain many of them side by side. The role You own models and data end to end: problem framing, features, training, evaluation, deployment, and the uncomfortable part where you explain to a customer why the vulnerability their board is worried about ranked 400th on the remediation list. What you'll do * Design, train, and ship exploit prediction models against ground-truth exploitation telemetry, in cloud, appsec, and traditional infrastructure. * Build evaluation that survives contact with reality. Precision, recall, coverage, efficiency, calibration, and how all four decay over time. Accuracy is not a number you report once at launch. * Solve for extreme class imbalance. A fraction of a percent of published CVEs are ever exploited in the wild, and most of the industry's modeling failures start with pretending that isn't true. * Work the hard part of the dual-model architecture: partial pooling, hierarchical priors, and cold-start behavior for customers whose local telemetry is thin in month one and rich in month twelve. * Engineer features across scanner output, EDR, asset inventory, identity, cloud posture, and exploitation telemetry, and be honest about which ones are leakage. * Own monitoring and drift detection. * Publish. Papers, methodology write-ups, open benchmarks, conference talks. Our positioning is that we show our work. * Partner with engineering and our forward deployed team to move models out of notebooks and into production systems that customers depend on. ## Related Videos - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [API Platform: From Rest & GraphQL APIs to state-of-the-art standards in seconds](https://www.wearedevelopers.com/videos/1968-api-platform-from-rest-graphql-apis-to-state-of-the-art-standards-in-seconds) - [Fault Tolerance and Consistency at Scale: Harnessing the Power of Distributed SQL Databases](https://www.wearedevelopers.com/videos/1146-fault-tolerance-and-consistency-at-scale-harnessing-the-power-of-distributed-sql-databases) - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [What non-automotive Machine Learning projects can learn from automotive Machine Learning projects](https://www.wearedevelopers.com/videos/397-what-non-automotive-machine-learning-projects-can-learn-from-automotive-machine-learning-projects) - [Machine Learning for Software Developers (and Knitters)](https://www.wearedevelopers.com/videos/154-machine-learning-for-software-developers-and-knitters) ## Related Articles - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers) - [Events like RSAC Get You CISOs. Developers Decide What Actually Gets Deployed.](https://www.wearedevelopers.com/magazine/693-events-like-rsac-get-you-cisos-developers-decide-what-actually-gets-deployed) - [Trustworthy AI Starts at Deployment: 5 Checks Before You Ship](https://www.wearedevelopers.com/magazine/753-trustworthy-ai-starts-at-deployment-5-checks-before-you-ship) - [Dev Digest 164: AI Agents, AI Blindspots and MCP security problems](https://www.wearedevelopers.com/magazine/578-dev-digest-164-ai-agents-ai-blindspots-and-mcp-security-problems)