{"slug": "guardianbench-a-same-scene-instruction-contrastive-benchmark-for-latent-risk-in", "title": "GuardianBench: A Same-Scene Instruction-Contrastive Benchmark for Latent Contextual Risk in Embodied AI", "summary": "Researchers introduced GuardianBench, a benchmark with 3,024 instruction-scene examples designed to test latent contextual risk in embodied AI, finding that state-of-the-art vision-language models achieve only 24.1% average pair accuracy. The study, posted on arXiv (2608.21928v1), identifies instruction-insensitive verdicts as the primary failure mode and proposes Verdict Log-Odds Supervision (VLOS) to improve performance on open-weight backbones.", "body_md": "arXiv:2608.21928v1 Announce Type: new\nAbstract: In embodied AI, safety risk can be latent: a benign instruction and a safe scene become hazardous only when composed. Prior work has advanced embodied safety by varying visual contexts or evaluating execution-time dynamics, but the complementary axis of fixing the scene and varying only the instruction remains underexplored. We introduce GuardianBench, an instruction-contrastive benchmark grounded in international safety standards that isolates this latent contextual risk through 3,024 instruction-scene examples organized as same-scene Safe/Unsafe contrastive pairs across various hazard categories. Benchmarking state-of-the-art vision-language models (VLMs) reveals instruction-insensitive verdicts: models disproportionately approve both instructions under a given scene; across the primary models, average pair accuracy is only 24.1%. Our systematic rationale audit localizes the dominant failure: models fail to bind the instruction-relevant cues that differentiate safe from unsafe compositions. As a post-training case study, Verdict Log-Odds Supervision (VLOS), a lightweight verdict-level objective, substantially improves performance on open-weight backbones. Together, our latent contextual risk task formulation, standards-grounded contrastive benchmark construction, pair-level and rationale-level failure diagnosis, and benchmark-enabled verdict calibration establish GuardianBench as a controlled evaluation suite for exposing and improving safety reasoning over instruction-scene compositions under latent contextual risk.", "url": "https://wpnews.pro/news/guardianbench-a-same-scene-instruction-contrastive-benchmark-for-latent-risk-in", "canonical_source": "https://www.machinebrief.com/news/guardianbench-a-same-scene-instruction-contrastive-benchmark-203g", "published_at": "2026-08-25 04:00:00+00:00", "updated_at": "2026-08-25 05:14:12.468444+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-safety", "ai-research", "computer-vision"], "entities": ["GuardianBench", "arXiv", "Verdict Log-Odds Supervision"], "alternates": {"html": "https://wpnews.pro/news/guardianbench-a-same-scene-instruction-contrastive-benchmark-for-latent-risk-in", "markdown": "https://wpnews.pro/news/guardianbench-a-same-scene-instruction-contrastive-benchmark-for-latent-risk-in.md", "text": "https://wpnews.pro/news/guardianbench-a-same-scene-instruction-contrastive-benchmark-for-latent-risk-in.txt", "jsonld": "https://wpnews.pro/news/guardianbench-a-same-scene-instruction-contrastive-benchmark-for-latent-risk-in.jsonld"}}