Your AI Can Say “Gravity” Without Knowing What It Means
A new analysis of vision-language models (VLMs) shows they fail basic physics reasoning, with the average model scoring around 40% on the PhysBench benchmark (ICLR 2025) and the best model, GPT-4o, scoring 49.49%, compar…