AI Breakdown
AI Breakdown
agibreakdown
ICCV 2023 - Sigmoid Loss for Language Image Pre-Training
3 minutes Posted Oct 17, 2023 at 3:48 pm.
0:00
3:30
Download MP3
Show notes
In this episode we discuss Sigmoid Loss for Language Image Pre-Training
by Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas Beyer. The paper introduces a pairwise Sigmoid loss for Language-Image Pre-training (SigLIP), which operates on image-text pairs and allows for scaling up batch size without the need for global pairwise similarities. By combining SigLIP with Locked-image Tuning, the authors achieve high ImageNet zero-shot accuracy in just two days of training. The authors also discuss the impact of batch size and find that a batch size of 32k is sufficient.