When Modality Gap Reduction Fails: Prediction-Level Hubness in CLIP

#When #Modality #Reduction #Fails #Prediction-Level

arXiv:2609.01103v1 Announce Type: new Abstract: Reducing the modality gap between image and text representations in CLIP is widely expected to improve cross-modal alignment and downstream performance. However, a smaller average image-text gap does not necessarily lead to consistent accuracy gains.…