Data Scientist II

Socure Inc.
San Francisco, CA, United States
1 day ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
1 year minimum
Working hours
Regular working hours
Job source

Tech stack

Airflow Amazon Web Services Amazon S3 Data Analysis Big Data Unix Code Review Extract Transform Load (ETL) Data Mining Distributed Systems Amazon DynamoDB Elasticsearch
+18 more
Graph Database Python (Programming Language) Machine Learning Neo4j Tensorflow SQL Databases Workflow Management Systems Data Processing Feature Engineering Sql Optimization Pytorch Apache Spark Data Lakes Pyspark Scikit Learn Xgboost Data Pipelines Databricks

Job description

The Big Data R&D team is responsible for building the core identity graph and entity-resolution capabilities that power Socure’s Deceased Monitoring and compliance products. In this role, you will help develop graph-based algorithms and data pipelines on massive PII datasets, support modelers with high-quality features, and evaluate new data sources that feed our identity and fraud products. You will work closely with senior data scientists and engineers while developing your skills in large-scale ML, distributed systems, and graph analytics.

What You’ll Do

  • Contribute to the design and implementation of machine learning, data mining, statistical, and graph-based algorithms to analyze very large datasets for identity verification and anomaly detection.
  • Analyze large datasets to help develop and refine entity-resolution and identity-matching algorithms that drive Socure’s Deceased Monitoring and compliance solutions.
  • Build and maintain components of data-processing pipelines (ETL, feature generation, normalization) using tools such as Spark/PySpark and AWS (e.g., EMR, S3).
  • Support senior data scientists with feature engineering, data exploration, error analysis, and A/B test setup for new models and signals.
  • Help evaluate new third-party and internal data sources: profile data quality, design offline experiments, and summarize impact on coverage and model performance.
  • Implement and maintain SQL and Python/R code for data extraction, transformation, and validation; contribute to code reviews and basic testing.
  • Provide analytical support to compliance and regulatory product teams, including ad hoc investigations, simple dashboards, and data deep dives.
  • Communicate findings in a clear, structured way to peers and cross-functional partners (Product, Engineering, Client Analysis), focusing on key insights and trade-offs.
  • Work effectively in a fast-paced, cross-functional environment; demonstrate ownership of well-scoped tasks and follow through to completion.

Requirements

  • Master’s degree with 2+ years of experience, or Ph.D. with 1+ years of experience in a data science or analytics role, or equivalent practical experience.
  • Proficiency in at least one general-purpose programming language used in data science (Python, or Scala).
  • Solid experience writing and optimizing SQL for large datasets; comfort working in data lake / warehouse environments.
  • Hands-on experience with Spark or PySpark and common ML libraries (e.g., scikit-learn, XGBoost, TensorFlow/PyTorch a plus).
  • Familiarity with UNIX environments and the AWS ecosystem (e.g., EMR, S3); Databricks experience is a plus.
  • Working knowledge of supervised/unsupervised ML and basic statistics (similarity measures, clustering, evaluation metrics).
  • Exposure to graph techniques or graph databases (Neo4j, AWS Neptune, GraphFrames) is a strong plus.
  • Bonus: experience with Elasticsearch or DynamoDB; workflow tools such as Airflow for automating data pipelines.
  • Ability to break down loosely defined problems, ask good clarifying questions, and iterate quickly with feedback.

Please note that sponsorship is not available at this time; and that you must be located within 45 miles of a talent hub to be considered.

About the company

Socure is building the identity trust infrastructure for the digital economy - verifying 100% of good identities in real time and stopping fraud before it starts. The mission is big, the problems are complex, and the impact is felt by businesses, governments, and millions of people every day.

We hire people who want that level of responsibility. People who move fast, think critically, act like owners, and care deeply about solving customer problems with precision. If you want predictability or narrow scope, this won’t be your place. If you want to help build the future of identity with a team that holds a high bar for itself - keep reading.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:09 min

Configuring synthetic data for safe interactive programming

Mingshen Sun Mingshen Sun · World Congress 2024

2:24 min

Comparing Neo4j and GraphQL conceptual models

William Lyon · LIVE

2:03 min

Microsoft integrating native Unix coreutils into Windows environments

Chris Heilmann +2 · LIVE

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

9:47 min

Transforming tabular metrics into meaningful business value dashboards

Boris Krumrey +2 · LIVE

3:30 min

Introduction to Neo4j and remote developer relations work

Videos

See all

Related articles

See all