GuardianBench: A Same-Scene Instruction-Contrastive Benchmark for Latent Contextual Risk in Embodied AI
Researchers introduced GuardianBench, a benchmark with 3,024 instruction-scene examples designed to test latent contextual risk in embodied AI, finding that state-of-the-art vision-language models achieve only 24.1% aver…