WeAreDevelopers LIVE • Jun 12, 2020

Anomaly Detection - Using unsupervised Machine Learning for detecting anomalies in customer base

Lukas Kölbl

Faced with scattered, subjective data, an insurer abandoned supervised learning. See how they used robust PCA to build an explainable anomaly detection model that caught 84% of true outliers.

Pause
Mute Enter Fullscreen
#1 about 4 min

Identifying business value through targeted data science

How data scientists bridge the gap between large data pipelines and actionable business insights.

#2 about 2 min

Combining technical skills with strong business communication

Why translating mathematical foundations into understandable language drives successful analytics initiatives.

#3 about 3 min

Balancing model training with data preparation realities

How real-world data science requires significant effort in data cleansing and collaboration rather than just modeling.

#4 about 3 min

Framing the customer anomaly detection use case

Why an insurance company transitioned from manual customer investigations to automated outlier detection.

#5 about 1 min

Ensuring unbiased results and interpretable model explainability

How removing demographic characteristics from algorithms ensures holistic and fair analytical outcomes.

#6 about 3 min

Structuring the machine learning solution approach architecture

Why moving from a supervised to an unsupervised approach successfully mitigates unstructured label data challenges.

#7 about 2 min

Building the analytical record for customer master data

How spending majority project time harmonizing datasets enables high-quality feature engineering and holistic views.

#8 about 2 min

Selecting and iterating on anomaly detection algorithms

Continuous iteration loops with business units help evaluate approaches like isolation forests and autoencoders.

#9 about 3 min

Implementing robust principal component analysis for outliers

Applying modified linear dimensionality reduction techniques effectively separates abnormal behavior from standard peer groups.

#10 about 1 min

Extracting explainability from orthogonal distance metrics

Using reconstruction errors easily identifies which specific feature triggers an outlier classification without complex libraries.

#11 about 2 min

Evaluating production model performance using a confusion matrix

Deploying the automated engine demonstrates high accuracy in classifying standard workflows versus genuine operational outliers.

#12 about 3 min

Measuring business efficiency gains from deployed machine learning

Visualizing peer group behavior clarifies decision boundaries and significantly reduces the manual investigation time previously required.

Matching moments

1:54 min

Evaluating unsupervised anomaly detection model performance in banking

Stefan Donsa Stefan Donsa +1 · LIVE

1:41 min

Implementing semi-automated anomaly detection with human oversight

Katja Träumner

4:45 min

Applying dimensionality reduction for customer profile reconstruction

Stefan Donsa Stefan Donsa +1 · LIVE

2:59 min

Understanding anomaly detection models and reducing false positives

Raz Cohen · LIVE

1:36 min

Performing exploratory data analysis to uncover underlying patterns

Julian Joseph · LIVE

5:21 min

Preventing customer churn with predictive machine learning models

Andreas Christian · LIVE