In this talk, Pengyuan Li will present the team's efforts in building the Granite Vision model—a lightweight, large vision-language model specifically designed to excel in enterprise use cases, particularly in visual document understanding. Compared to other models of similar parameter size, Granite Vision achieves state-of-the-art performance on established benchmarks for visual document understanding, table and chart conversion to structured formats, and general image tasks. Since its release, Granite Vision has yield more than 150K usages and downloads on Hugging Face. The model is publicly available on HuggingFace under the Apache 2.0 license.
Become a paid member of the channel to help us make more episodes https://www.youtube.com/channel/UCkrcW82Y2kbgU-U9RaYfgxw/join
OpenCV is a 501(c)(3) registered non-profit in the United States. See how you can support open source CV & AI: http://opencv.org/support/

