Vision Language Grounding: How AI Connects “Dog” to Pixels, and Where It Falls Apart
OpenAI researchers demonstrated in 2021 that CLIP, a vision language model, misclassifies a Granny Smith apple as an iPod when a paper label reading 'iPod' is attached, revealing that grounding in AI …