> Markdown version of [/videos/301-making-neural-networks-portable-with-onnx](https://www.wearedevelopers.com/videos/301-making-neural-networks-portable-with-onnx). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Making neural networks portable with ONNX Struggling to deploy Python models in Java or JavaScript? ONNX acts as a universal translator for neural networks. Decouple training from deployment and ship AI to the edge seamlessly. - **Speakers:** Ron Dagdag - **Event:** JavaScript Congress - **Published:** November 25, 2021 - **Duration:** 54:49 - **URL:** https://www.wearedevelopers.com/videos/301-making-neural-networks-portable-with-onnx ## Summary Machine learning teams often struggle to deploy models trained in Python-based frameworks like PyTorch or TensorFlow into production environments running C#, Java, or JavaScript. The core solution is the Open Neural Network Exchange (ONNX), an open and portable format for machine learning models that acts as a universal intermediate representation, much like a PDF does for text documents. By decoupling the model's training framework from its production deployment, ONNX bridges the gap between the data scientist generating the algorithm and the software engineer scaling it for real-world applications. Creating or acquiring these models is highly flexible; teams can convert existing models, train new ones in the cloud, export from Custom Vision services, or instantly download pre-trained state-of-the-art architectures directly from the ONNX Model Zoo.\n\nOnce an ONNX model is ready, developers can use visualization tools like Netron to inspect its graph of operations, establishing a clear understanding of the matrix dimensions required for data inputs and outputs. Execution is then handled by ONNX Runtime, a high-performance inference engine capable of running in the cloud, on localized fog networks, or directly on edge devices. Deploying AI workloads to the edge—whether via a Node.js backend using WebAssembly or directly within a web browser via WebGL—drastically reduces inferencing latency while improving data privacy, offline availability, and cloud ingestion costs.\n\nUltimately, successfully deploying machine learning models dictates their real-world value. Data scientists function much like master bakers experimenting with ingredients to perfect a recipe, but software engineers run the bakery itself—positioning the product, minimizing supply chain bottlenecks, and ensuring reliable customer delivery. By leveraging ONNX and its runtime ecosystem, development teams can seamlessly integrate complex AI logic into standard web and mobile architectures, turning specialized academic models into viable, cross-platform business solutions. **Keywords:** onnx format, neural network portability, onnx runtime, machine learning inferencing, python model conversion, onnx model zoo, edge deployment latency, webassembly machine learning, hardware acceleration compatibility, netron visualization, javascript ml integration, custom vision models, cross-framework deployment, edge computing privacy, node.js machine learning ## Chapters 1. **Establishing introductory context regarding artificial intelligence deployments** (00:00) — Framing context around artificial intelligence deployment capabilities establishes foundational concepts for developers transitioning operational logic structures. 1. **Differences between traditional programming and machine learning algorithms** (02:01) — Providing structured data as training examples allows programmatic pipelines to automatically write an algorithm solving a core problem. 1. **Bridging application frameworks with the universal ONNX format** (04:39) — Creating universal model formats allows artificial intelligence capabilities to run seamlessly across completely disconnected server architectures. 1. **Ideal system architectures and use cases for portability** (07:59) — Removing heavy dependencies reduces computational latency making models ideal for constrained environments running directly on hardware endpoints. 1. **Sourcing and generating pre-trained machine learning models** (10:44) — Bootstrapping software capabilities via public model repositories accelerates immediate project value avoiding prolonged dataset configuration overhead. 1. **Visualizing model operation graphs with the Netron application** (14:51) — Inspecting loaded payload architectures illuminates proper array dimensions and variable names needed when coding inference execution scripts. 1. **Converting existing application architectures into interoperable runtime graphs** (16:29) — Recompiling localized scripts into universally accessible payloads unifies continuous integration infrastructure when managing experimental operational variants. 1. **Comparing cloud compute infrastructure against local edge deployment** (22:06) — Distributing computation targets closer to client networks lowers cloud ingestion fees and resolves critical bandwidth dependency factors. 1. **Integrating capabilities through high performance ONNX runtime engines** (28:19) — Abstracting execution paths ensures analytical computations fall gracefully back on existing hardware graphic processors determining automatic render speeds. 1. **Writing data payload mappings inside Node backend instances** (29:56) — Translating primitive numbers into valid session instances enables raw backend environments securely interpreting loaded runtime regression logic definitions. 1. **Running independent inference graphs securely on client browsers** (33:30) — Isolating execution evaluations natively within standard browser sessions protects client transmission events and avoids continuous architectural hosting limits. 1. **Processing browser image states into acceptable neural formats** (36:05) — Programmatically resizing source images matching specified training dimensions accurately transforms raw inputs correctly predicting contextual emotion markers. 1. **Optimizing graph layers targeting constrained mobile execution environments** (43:02) — Optimizing parameter weight scales tightens application boundaries preventing bloated bundle downloads restricting underlying runtime mobile execution capabilities. 1. **Addressing audience questions on entering data science domains** (45:00) — Exploring straightforward remote interfaces builds comfortable developer understanding before studying complex foundational computation architecture initially required completely. ## Related Moments - [Standardizing interoperable model deployments with the ONNX framework](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) (from "Introduction to Azure Machine Learning") - [Exporting framework independent representations for edge processing platforms](https://www.wearedevelopers.com/videos/367-intelligent-data-selection-for-continual-learning-of-ai-functions) (from "Intelligent Data Selection for Continual Learning of AI Functions") - [Open-source community and machine learning frameworks](https://www.wearedevelopers.com/videos/1420-mobile-ai-just-got-faster-what-s-coming-for-developers-on-arm) (from "Mobile AI Just Got Faster: What’s Coming for Developers on Arm") - [Leveraging ONNX Runtime Web for local model execution](https://www.wearedevelopers.com/videos/1572-privacy-first-in-browser-generative-ai-web-apps-offline-ready-future-proof-standards-based) (from "Privacy-first in-browser Generative AI web apps: offline-ready, future-proof, standards-based") - [Exporting and evaluating learned ONNX models](https://www.wearedevelopers.com/videos/272-machine-learning-in-ml-net) (from "Machine Learning in ML.NET") - [Comparing ONNX runtime web and TensorFlow deployments](https://www.wearedevelopers.com/videos/1896-how-web-ai-can-power-the-agentic-web-jason-mayes-google) (from "How Web AI Can Power the Agentic Web - Jason Mayes (Google)") ## Related Articles - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) ## Related Jobs - [AI Software Engineer (Germany)](https://www.wearedevelopers.com/jobs/48317-ai-software-engineer-germany) at **Sunhat** - [Machine Learning Engineer](https://www.wearedevelopers.com/jobs/ext/588393-machine-learning-engineer) at **Twilio** - [Principal Engineer - AI Search & Vector Infrastructure](https://www.wearedevelopers.com/jobs/ext/353953-principal-engineer-ai-search-vector-infrastructure) at **Redis** - [Data Scientist](https://www.wearedevelopers.com/jobs/ext/1351648-data-scientist) at **Almedia** - [Machine Learning Engineer](https://www.wearedevelopers.com/jobs/ext/1355348-machine-learning-engineer) at **TWILIO** - [Principal Software Engineer, Enterprise AI Platform](https://www.wearedevelopers.com/jobs/ext/1467292-principal-software-engineer-enterprise-ai-platform) at **GitHub**