16:24
2026-09-11
promptcube3.com
large-language-models
Why MCQ-based evaluation beats BLEU for video captions
A proposed evaluation method for video captioning replaces BLEU and METEOR with Multiple-Choice Question Answering (MCQA), scoring captions by the percentage of video-grounded questions an LLM can ansβ¦