16:31
2026-08-10
pub.towardsai.net
artificial-intelligence
Your AI Can Say βGravityβ Without Knowing What It Means
A new analysis of vision-language models (VLMs) shows they fail basic physics reasoning, with the average model scoring around 40% on the PhysBench benchmark (ICLR 2025) and the best model, GPT-4o, scβ¦