> Markdown version of [/videos/6-anomaly-detection-using-unsupervised-machine-learning-for-detecting-anomalies-in-customer-base?t=212](https://www.wearedevelopers.com/videos/6-anomaly-detection-using-unsupervised-machine-learning-for-detecting-anomalies-in-customer-base?t=212). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Anomaly Detection - Using unsupervised Machine Learning for detecting anomalies in customer base Faced with scattered, subjective data, an insurer abandoned supervised learning. See how they used robust PCA to build an explainable anomaly detection model that caught 84% of true outliers. - **Speakers:** Lukas Kölbl - **Event:** WeAreDevelopers LIVE - **Published:** June 12, 2020 - **Duration:** 22:06 - **URL:** https://www.wearedevelopers.com/videos/6-anomaly-detection-using-unsupervised-machine-learning-for-detecting-anomalies-in-customer-base ## Summary Effective data science requires deep technical expertise combined with strong business alignment. While the industry often fixates on model training, the reality is that data cleansing, feature engineering, and understanding the core business problem consume the vast majority of project lifecycles. Constructing a cohesive analytical record by harmonizing fragmented customer datasets typically requires about 70% of total project effort. This solid data foundation is a prerequisite for moving beyond siloed, manual workflows and successfully tackling complex business challenges. In a practical anomaly detection application for a European insurance provider, replacing manual customer profiling with an automated framework drastically improved operational efficiency. The initial goal was to build a product recommendation engine, but an unbiased outlier detection approach proved far more effective for improving customer satisfaction. Because the organization's existing labeled data was subjective and scattered across inconsistent spreadsheets, the strategy pivoted away from supervised learning toward unsupervised machine learning. Focusing on holistic customer behaviors rather than specific demographic traits ensured the final model remained unbiased and non-discriminatory. Evaluating several algorithms, including Isolation Forests and Autoencoders, led to the selection of a robust Principal Component Analysis (PCA) methodology. Relying on dimensionality reduction, this framework identifies anomalies based on reconstruction error, calculating orthogonal and score distances to judge if a particular profile matches a standard peer group. A major advantage of using a linear model like robust PCA is its inherent explainability, which provides direct insights into why a user was flagged without requiring heavy interpretability tools like LIME or SHAP. Deployed into production alongside a clear traffic-light visualization system, the model successfully detected 84% of true outliers, vastly outperforming manual human evaluation in both accuracy and speed. **Keywords:** anomaly detection methods, unsupervised machine learning, customer base outliers, business alignment in data science, data harmonization techniques, analytical record creation, robust PCA applications, feature engineering workflows, dimensionality reduction algorithms, data reconstruction error, orthogonal distance measurement, score distance evaluation, machine learning explainability, insurance customer analytics, unbiased predictive modeling, false positive reduction ## Chapters 1. **Identifying business value through targeted data science** (00:16) — How data scientists bridge the gap between large data pipelines and actionable business insights. 1. **Combining technical skills with strong business communication** (03:32) — Why translating mathematical foundations into understandable language drives successful analytics initiatives. 1. **Balancing model training with data preparation realities** (05:15) — How real-world data science requires significant effort in data cleansing and collaboration rather than just modeling. 1. **Framing the customer anomaly detection use case** (07:33) — Why an insurance company transitioned from manual customer investigations to automated outlier detection. 1. **Ensuring unbiased results and interpretable model explainability** (09:45) — How removing demographic characteristics from algorithms ensures holistic and fair analytical outcomes. 1. **Structuring the machine learning solution approach architecture** (10:44) — Why moving from a supervised to an unsupervised approach successfully mitigates unstructured label data challenges. 1. **Building the analytical record for customer master data** (12:47) — How spending majority project time harmonizing datasets enables high-quality feature engineering and holistic views. 1. **Selecting and iterating on anomaly detection algorithms** (14:06) — Continuous iteration loops with business units help evaluate approaches like isolation forests and autoencoders. 1. **Implementing robust principal component analysis for outliers** (15:10) — Applying modified linear dimensionality reduction techniques effectively separates abnormal behavior from standard peer groups. 1. **Extracting explainability from orthogonal distance metrics** (17:30) — Using reconstruction errors easily identifies which specific feature triggers an outlier classification without complex libraries. 1. **Evaluating production model performance using a confusion matrix** (18:04) — Deploying the automated engine demonstrates high accuracy in classifying standard workflows versus genuine operational outliers. 1. **Measuring business efficiency gains from deployed machine learning** (19:50) — Visualizing peer group behavior clarifies decision boundaries and significantly reduces the manual investigation time previously required. ## Related Moments - [Evaluating unsupervised anomaly detection model performance in banking](https://www.wearedevelopers.com/videos/111-detecting-money-laundering-with-ai) (from "Detecting Money Laundering with AI") - [Implementing semi-automated anomaly detection with human oversight](https://www.wearedevelopers.com/videos/1308-data-science-ml-ai-in-the-oil-and-gas-industry-at-ndt-global-dr-katja-traumner) (from "Data Science, ML & AI in the Oil and Gas Industry at NDT Global - Dr. Katja Träumner") - [Applying dimensionality reduction for customer profile reconstruction](https://www.wearedevelopers.com/videos/111-detecting-money-laundering-with-ai) (from "Detecting Money Laundering with AI") - [Understanding anomaly detection models and reducing false positives](https://www.wearedevelopers.com/videos/853-navigating-the-ai-wave-in-devops) (from "Navigating the AI Wave in DevOps") - [Performing exploratory data analysis to uncover underlying patterns](https://www.wearedevelopers.com/videos/586-data-science-in-retail) (from "Data Science in Retail") - [Preventing customer churn with predictive machine learning models](https://www.wearedevelopers.com/videos/253-data-fabric-in-action-how-to-enhance-a-stock-trading-app-with-ml-and-data-virtualization) (from "Data Fabric in Action - How to enhance a Stock Trading App with ML and Data Virtualization") ## Related Articles - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [How machine learning can help us tell fact from fiction](https://www.wearedevelopers.com/magazine/509-how-machine-learning-can-help-us-tell-fact-from-fiction) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) ## Related Jobs - [Data Scientist](https://www.wearedevelopers.com/jobs/ext/1351648-data-scientist) at **Almedia** - [Principal Machine Learning Engineer](https://www.wearedevelopers.com/jobs/ext/1706410-principal-machine-learning-engineer) at **Almedia** - [Machine Learning Engineer](https://www.wearedevelopers.com/jobs/ext/1377841-machine-learning-engineer) at **Almedia** - [Data & Machine Learning Engineer | Hybrid work](https://www.wearedevelopers.com/jobs/ext/431779-data-machine-learning-engineer-hybrid-work) at **SMG Swiss Marketplace Group** - [AI Software Engineer (Germany)](https://www.wearedevelopers.com/jobs/48317-ai-software-engineer-germany) at **Sunhat** - [Senior Software Engineer, Data](https://www.wearedevelopers.com/jobs/48273-senior-software-engineer-data) at **Sportradar Media Services GmbH**