Weaviate Podcast
Weaviate Podcast
Weaviate
Unstructured with Brian Raymond - Weaviate Podcast #48!
43 minutes Posted May 23, 2023 at 2:00 pm.
Welcome Brian!!
What is Unstructured?
Why now? New Advancements in Unstructured
Thoughts on Data Connectors Hub
PDFs to Weaviate with Unstructured
State-of-the-Art in OCR and Document Parsing
How to get the data from Weaviate.io?
Foundation Models from Unstructured
Evaporate-Code+
CSV, Parquet, JSON transformations in Staging
Cleaning Bricks
Visual Document Examples
Text Chunking with Metadata
Knowledge Graphs with Goldman Sachs example
LLM Hallucinations
Announcements from Brian!
0:00
43:03
Download MP3
Show notes
Hey everyone, thank you so much for watching the 48th episode of the Weaviate Podcast!! This is a SUPER exciting one, welcoming Brian Raymond the CEO / Founder of Unstructured! Unstructured is a perfect complimenting technology for Weaviate, helping people get their Unstructured data into Weaviate! The podcast dives into the nuances of this task, but it generally revolves around Unstructured's abstraction of Partitioning, Cleaning, and Staging! Unstructured is making groundbreaking innovations on using Visual Document Layout models for Partitioning, for example saying that this part of the PDF is the header, body, image caption, and so on. Cleaning then describes removing pesky details like whitespaces or odd characters. Staging then describes the transformations of say formatting a text chunk with it's metadata into the JSON for a Weaviate object upload! I really hope you find this podcast interesting! We are publishing a blog post as well showing an example of how to use Unstructured to get PDF data into Weaviate, please please check that out and let us know if it works for your data and how we can improve it! This blog post can be found on weaviate.io and we will be managing discussions around it both in the Weaviate slack, as well as Unstructured! Thank you so much for listening!
Check out Unstructured here! https://www.unstructured.io/
Chapters