PROBE: Manipulation-Grounded Visual Question Answering with VLM Agents
Researchers introduced PROBE, a framework for benchmarking and finetuning vision-language model (VLM) agents on Manipulation-Grounded Visual Question Answering (MG-VQA), where robots must physically move objects to answe…