> Markdown version of [/videos/857-convert-batch-code-into-streaming-with-python?t=2088](https://www.wearedevelopers.com/videos/857-convert-batch-code-into-streaming-with-python?t=2088). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Convert batch code into streaming with Python What if you could switch from static batch data to live streaming with one configuration change? Learn to ditch complex JVMs and power real-time AI applications using pure Python. - **Speakers:** Bobur Umurzokov - **Event:** WeAreDevelopers LIVE - **Published:** February 23, 2024 - **Duration:** 46:43 - **URL:** https://www.wearedevelopers.com/videos/857-convert-batch-code-into-streaming-with-python ## Summary Traditionally, building streaming pipelines required juggling complex, separate ecosystems for data ingestion and processing—often relying on Java, JVMs, and Kafka clusters. This video introduces a simplified approach using Python-based frameworks, spotlighting Pathway. By unifying the streaming data platform and processor, developers can focus purely on business logic without learning new infrastructure. Crucially, Pathway allows engineers to build and test pipelines on static batch data, then convert them to live streaming mode with a single configuration change, bypassing the notorious integration friction that stalls migrations. Beyond standard data engineering, this real-time paradigm unlocks powerful capabilities for AI applications. Generative AI tools, particularly Retrieval-Augmented Generation (RAG) systems, frequently suffer from stale context and cumbersome vector database synchronization. Real-time frameworks circumvent this by performing dynamic data indexing directly in memory. As source documents or live APIs change, updates are instantly detected, re-chunked, and embedded. This allows LLMs to query the absolute latest information without manual syncing or relying on complex external orchestration layers. The session also unpacks the mechanics of complex event processing, highlighting how dynamic table schemas automatically adapt to new columns without requiring pipeline restarts or table drops. By embedding real-time stream processing directly into Python applications, teams can reduce cloud compute bills compared to legacy solutions like Apache Spark or Flink. Ultimately, adopting a unified streaming approach accelerates the deployment of event-driven microservices and ensures AI agents operate on live, accurate context. **Keywords:** python stream processing, batch to streaming migration, pathway framework, real-time data pipelines, RAG architecture, vector database synchronization, real-time data indexing, event-driven microservices, LLM application development, change data capture, in-memory data processing, dynamic schema updates, apache flink alternatives, developer advocacy ## Chapters 1. **Advantages of Python frameworks for data streaming** (00:02) — An introduction to Python's capability for unifying stream processing and data platforms without complex management layers. 1. **Use cases for stream processing with Python frameworks** (05:31) — Python stream processing complements microservices architectures and enables real-time vector embeddings for machine learning. 1. **Building unified processing pipelines with the Pathway framework** (07:28) — Pathway provides a single python code base to smoothly transition from testing on static data to running streaming inputs. 1. **Integrating real-time data indexing into AI applications** (12:35) — Real-time continuous updates generate vector embeddings for augmented generation without requiring separate vector storage synchronizations. 1. **Examples of real-time artificial intelligence application development** (15:47) — Practical applications include summarizing unstructured documents, identifying real-time market discounts, and generating automated alerts. 1. **Code and architecture for building the discount application** (19:26) — A walkthrough of a python backend that extracts data and utilizes user interfaces via streamlit and openai endpoints. 1. **Demonstrating a real-time Dropbox document summarization application** (22:36) — A dockerized application chunks localized text and continuously polls for query responses to summarize expense invoices. 1. **Final takeaways on Python for streaming and batch processing** (27:38) — Python frameworks simplify the transition to streaming workloads by natively compiling python logic without relying on java abstractions. 1. **Understanding how Pathway ensures low latency data processing** (29:22) — Dynamic detection of newly added table row streams enables rapid state updates directly within memory-based inputs. 1. **Handling complex event processing and pattern recognition formats** (31:09) — Kafka event brokers seamlessly ingest dynamic column adjustments within tables without entirely dropping schema configurations. 1. **Limitations when migrating batch processes to native streaming** (33:04) — Migrating to streaming infrastructures introduces common limitations regarding learning curves, connector incompatibility, and complete pipeline rebuilds. 1. **Managing automatic load balancing and data skew effectively** (34:48) — Managed services balance processing loads by splitting transformations across instances and continuously merging output states. 1. **Implementing data parallelism constraints to improve processing speeds** (35:48) — Breaking complex transformation logic into smaller steps guarantees optimal scaling mechanisms without blocking sequential dependencies. 1. **Impact of real-time processing on organizational decision making** (37:10) — Continually cycling analytics enables faster customer experiences by prioritizing fresh data visualizations over overnight batch delays. 1. **Future evolution of real-time data processing technologies** (38:48) — Startups are transitioning storage tools into real-time streaming databases that combine unified architectures natively. 1. **Resource utilization techniques in real-time processing environments** (41:34) — Restricting continuous memory operations eliminates the excessive data cluster expansion and storage distributions required by rigid frameworks. 1. **Improving user experiences using directly connected real-time sources** (43:03) — Removing backend logic layers empowers frontline components to autonomously generate dynamic visualizations driven by real-time streams. 1. **Defining the core responsibilities of a developer advocate** (44:57) — Developer advocates leverage their technical software backgrounds to effectively communicate product benefits directly with engineering teams. ## Related Moments - [Rise of Python in real-time data processing](https://www.wearedevelopers.com/videos/1233-python-based-data-streaming-pipelines-within-minutes) (from "Python-Based Data Streaming Pipelines Within Minutes") - [Overcoming typical barriers to real-time stream processing adoption](https://www.wearedevelopers.com/videos/1233-python-based-data-streaming-pipelines-within-minutes) (from "Python-Based Data Streaming Pipelines Within Minutes") - [Introducing data management and the shift to streaming](https://www.wearedevelopers.com/videos/538-event-messaging-and-streaming-with-apache-pulsar) (from "Event Messaging and Streaming with Apache Pulsar") - [Comparing stream processing architecture to traditional batch processing](https://www.wearedevelopers.com/videos/1233-python-based-data-streaming-pipelines-within-minutes) (from "Python-Based Data Streaming Pipelines Within Minutes") - [Recognizing architectural drivers pushing event streaming system adoption](https://www.wearedevelopers.com/videos/538-event-messaging-and-streaming-with-apache-pulsar) (from "Event Messaging and Streaming with Apache Pulsar") - [Summarizing Python frameworks advantages and future event streams](https://www.wearedevelopers.com/videos/1233-python-based-data-streaming-pipelines-within-minutes) (from "Python-Based Data Streaming Pipelines Within Minutes") ## Related Articles - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [Why Event-Driven Architecture Isn’t About Speed (and When You Actually Need It)](https://www.wearedevelopers.com/magazine/745-why-event-driven-architecture-isn-t-about-speed-and-when-you-actually-need-it) ## Related Jobs - [Software Engineer, Fullstack](https://www.wearedevelopers.com/jobs/48415-software-engineer-fullstack) at **Sciforium** - [Software Engineer, Python (Asset Pricing & Hedging)](https://www.wearedevelopers.com/jobs/ext/2772280-software-engineer-python-asset-pricing-hedging) at **Bitpanda** - [Senior AI Developer](https://www.wearedevelopers.com/jobs/ext/2836034-senior-ai-developer) at **PwC** - [LLM Training Engineer](https://www.wearedevelopers.com/jobs/48420-llm-training-engineer) at **Sciforium** - [Principal Software Engineer, AI Compute Platform](https://www.wearedevelopers.com/jobs/ext/2847710-principal-software-engineer-ai-compute-platform) at **ARM** - [Senior AI Serving Engineer, Backend](https://www.wearedevelopers.com/jobs/48414-senior-ai-serving-engineer-backend) at **Sciforium**