Weaviate Podcast
Weaviate Podcast
Weaviate
Rohit Agarwal on Portkey - Weaviate Podcast #61!
49 minutes Posted Aug 3, 2023 at 2:00 pm.
Introduction
Portkey, Founding Vision
LLMOps vs. MLOps
Inference Hosting Options
3 Layers of LLM Use
LLM Load Balancers
Fine-Tuning LLMs
Retrieval-Aware Tuning
Portkey Cost Savings
HuggingGPT
Semantic Caching
Frequently Asked Questions
Embeddings vs. Generative Tasks
AI Moats, GPT Wrappers
Unlocks from Cheaper LLM Inference
0:00
49:24
Download MP3
Show notes
Hey everyone! Thank you so much for watching the 61st episode of the Weaviate Podcast! I am beyond excited to publish this one! I first met Rohit at the Cal Hacks event hosted by UC Berkeley where we had a debate about the impact of Semantic Caching! Rohit taught me a ton about the topic and I think it's going to be one of the most impactful early applications of Generative Feedback Loops! Rohit is building Portkey, a SUPER interesting LLM middleware that does things like load balancing between LLM APIs, and as discussed in the podcast there are all sorts of opportunities for this kind of space whether it be routing to tool-specific LLMs, different cost / accuracy requirements, or multiple models in the HuggingGPT sense. It was amazing chatting with Rohit, this was the best dive into LLMOps I have personally been apart of! As always we are more than happy to answer any questions or discuss any ideas you have about the content in the podcast!
Check out portkey here! https://portkey.ai/blog
Chapters