> Markdown version of [/videos/396-what-is-relational-learning-and-why-does-it-matter](https://www.wearedevelopers.com/videos/396-what-is-relational-learning-and-why-does-it-matter). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # What is relational learning and why does it matter? Are you still wasting weeks manually engineering features for complex enterprise databases? Relational learning automates this bottleneck using statistical optimization to unlock rapid, end-to-end predictive analytics. - **Speakers:** Alexander Uhlig - **Event:** World Congress 2022 - **Published:** June 15, 2022 - **Duration:** 28:00 - **URL:** https://www.wearedevelopers.com/videos/396-what-is-relational-learning-and-why-does-it-matter ## Summary Machine learning models traditionally expect fixed-shape input vectors, which makes formatting structured images, text, or audio relatively straightforward through standard transformations. However, tackling real-world enterprise databases with unbounded one-to-many relationships—such as customer transaction histories—disrupts this paradigm. Attempting to fit complex relational data into fixed inputs usually forces data scientists to either truncate datasets, which sacrifices critical predictive information, or perform weeks of manual, trial-and-error feature engineering alongside domain experts. To automate this bottleneck, the industry historically leaned on "propositionalization," a brute-force approach that aggregates all possible statistical combinations using tools like Featuretools or tsfresh. While accessible, this method scales poorly on complex architectures like Pandas, running redundant queries and flooding the pipeline with unhelpful "garbage features." The modern alternative shifts strategy by applying supervised machine learning directly to the feature engineering phase. By utilizing statistical optimization rather than brute calculation, algorithms can dynamically search for the most relevant table merges and aggregate operations. Frameworks housing algorithms like Multirel, Relboost, and RelMT approach this by adapting multi-relational decision tree concepts. They rely on an iterative loss function to incrementally add constraints and evaluate aggregations, drastically shrinking the massive computational search space before finalizing the feature set. Often pairing an accessible Python API with a highly optimized C++ engine, relational learning allows engineering teams to deploy rapid, end-to-end predictive analytics directly over complex data schemas without getting bogged down in manual pipeline construction. **Keywords:** relational learning, automated feature engineering, propositionalization strategies, machine learning fixed-shape inputs, one-to-many database relationships, predictive analytics operations, manual feature extraction, garbage feature reduction, multirel algorithm, relboost framework, supervised dataset aggregation, multi-relational decision trees, c++ optimized data transformation, customer churn prediction, open-source featuretools ## Chapters 1. **Introduction to data transformations for machine learning** (00:15) — Most machine learning models require data scientists to perform simple transformations to create consistent input vectors. 1. **The challenge of predicting with relational data** (03:54) — Relational databases with one-to-many relationships resist standard preprocessing techniques and fixed input vector formatting. 1. **The limitations of manual feature engineering cycles** (06:42) — Relying on domain experts and manual aggregation leads to inefficient, week-long development cycles for complex predictions. 1. **Automating feature engineering for relational datasets** (09:11) — New algorithms can automatically learn features from relational data to provide end-to-end automatization of predictive analytics. 1. **Analyzing brute-force feature learning and propositionalization** (10:37) — Traditional brute-force approaches apply numerous aggregations to columns but are inefficient and produce unhelpful garbage features. 1. **Supervised search algorithms for feature optimization** (15:02) — Supervised learning techniques statistically optimize the search for the best merge and aggregate operations across datasets. 1. **Iterative feature learning using conditional multirel logic** (16:26) — Generalizing decision tree patterns to relational data iteratively adds conditions to directly improve a loss function. 1. **Python API integration for relational machine learning frameworks** (20:33) — Data scientists can use a familiar Python pipeline structure to efficiently define feature learners and predictors. 1. **Getting started with open-source fast-prop implementations** (21:15) — Options for immediate integration include existing tools or the highly efficient fast-prop implementation for rapid propositionalization. 1. **Deep neural networks versus relational data structures** (23:06) — Deep neural networks require fixed-length input vectors, making them incompatible with the unbound relationships of relational data. 1. **Iteratively updating loss functions across relational parameters** (24:04) — Iteratively updating the loss function evaluates the impact of minor condition variations across a massive feature space. 1. **Supervised search space versus brute-force calculation and pruning** (25:13) — Defining a localized search space around data joins proves much faster than calculating billions of feature combinations. ## Related Moments - [Audience questions on practical machine learning operational strategies](https://www.wearedevelopers.com/videos/262-is-my-ai-alive-but-brain-dead-how-monitoring-can-tell-you-if-your-machine-learning-stack-is-still-performing) (from "Is my AI alive but brain-dead? How monitoring can tell you if your machine learning stack is still performing") - [Revolutionizing database search queries with language models](https://www.wearedevelopers.com/videos/2088-plan-to-link-your-llm-to-your-production-database-what-could-possibly-go-wrong) (from "Plan to link your LLM to your production database? What could possibly go wrong?") - [Designing an automated and orchestrated machine learning target process](https://www.wearedevelopers.com/videos/185-effective-machine-learning-managing-complexity-with-mlops) (from "Effective Machine Learning - Managing Complexity with MLOps") - [Engineering scale-out architectures specifically for relational queries](https://www.wearedevelopers.com/videos/1146-fault-tolerance-and-consistency-at-scale-harnessing-the-power-of-distributed-sql-databases) (from "Fault Tolerance and Consistency at Scale: Harnessing the Power of Distributed SQL Databases") - [Storing machine learning data with specialized vector databases](https://www.wearedevelopers.com/videos/844-enter-the-brave-new-world-of-genai-with-vector-search) (from "Enter the Brave New World of GenAI with Vector Search") - [Choosing between rule-based queries and machine learning](https://www.wearedevelopers.com/videos/1430-logs-in-observability-correlation) (from "Logs in observability - Correlation") ## Related Articles - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) ## Related Jobs - [Machine Learning Engineer](https://www.wearedevelopers.com/jobs/ext/1355348-machine-learning-engineer) at **TWILIO** - [Machine Learning Engineer](https://www.wearedevelopers.com/jobs/ext/588393-machine-learning-engineer) at **Twilio** - [Data Scientist](https://www.wearedevelopers.com/jobs/ext/1351648-data-scientist) at **Almedia** - [Data & Machine Learning Engineer | Hybrid work](https://www.wearedevelopers.com/jobs/ext/431779-data-machine-learning-engineer-hybrid-work) at **SMG Swiss Marketplace Group** - [Principal Machine Learning Engineer](https://www.wearedevelopers.com/jobs/ext/1706410-principal-machine-learning-engineer) at **Almedia** - [Staff, Machine Learning Engineer (L4)](https://www.wearedevelopers.com/jobs/ext/1202639-staff-machine-learning-engineer-l4) at **Twilio**