WeAreDevelopers LIVE • Oct 12, 2020

PySpark - Combining Machine Learning & Big Data

Ayon Roy

Bypass complex Scala programming and deploy native Python operations across distributed clusters. Learn how PySpark and MLlib unite to accelerate robust, scalable machine learning pipelines.

Pause
Mute Enter Fullscreen
#1 about 4 min

Defining big data and machine learning fundamentals

The four Vs conceptualize big data while machine learning enables systems to predict outcomes without explicit programming.

#2 about 3 min

Why organizations combine big data and machine learning

Tapping into massive amounts of new data with machine learning drives personalized consumer experiences.

#3 about 3 min

Capabilities of the Apache Spark processing engine

Spark provides a large-scale data processing environment with integrated libraries for machine learning, SQL, and streaming.

#4 about 4 min

Understanding RDDs, DataFrames, and Datasets in Spark

Data ingestion evolved from complex resilient distributed datasets into developer-friendly structures like DataFrames and Datasets.

#5 about 5 min

Exploring the Apache Spark layer architecture

Cluster managers drive resource allocation while Spark Core processes diverse data sources using top-layer libraries.

#6 about 8 min

How cluster managers orchestrate distributed tasks

A driver program coordinates with a cluster manager via a Spark context to execute computations in parallel on worker nodes.

#7 about 5 min

Harnessing Spark with Python using PySpark and Py4J

The PySpark library relies on Py4J sockets to translate Python instructions into processes compatible with Java Virtual Machines.

#8 about 5 min

Advantages of utilizing Spark's native MLlib

Built-in classification, feature extraction, and pipeline support make MLlib a scalable choice for data analytics.

#9 about 5 min

Constructing ML pipelines with transformers and estimators

DataFrames pass through transformers to engineer features and enter estimators that output functional predictive models.

#10 about 4 min

Pre-built algorithms and further learning resources

Spark MLlib offers accessible, pre-built regression and clustering algorithms for integrating big data processing with predictive modeling.

Matching moments

2:56 min

Options for database machine learning integration architectures

Akmal Chaudhri Akmal Chaudhri · LIVE

4:18 min

Exploring the big data and machine learning portfolio

Qiyang Duan · LIVE

2:04 min

Comparing offline data analytics with online stream processing

Artem Volk Artem Volk +1 · World Congress 2024

1:41 min

Visualizing the complex developer journey for JVM ecosystems

Bobur Umurzokov · LIVE

15:08 min

Audience questions on practical machine learning operational strategies

Lina Weichbrodt · LIVE

7:10 min

Exploring pathways into the machine learning engineering field

Jose Luis Latorre Millas · LIVE