cd /news/artificial-intelligence/guardianbench-a-same-scene-instructi… · home topics artificial-intelligence article
[ARTICLE · art-109696] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

GuardianBench: A Same-Scene Instruction-Contrastive Benchmark for Latent Contextual Risk in Embodied AI

Researchers introduced GuardianBench, a benchmark with 3,024 instruction-scene examples designed to test latent contextual risk in embodied AI, finding that state-of-the-art vision-language models achieve only 24.1% average pair accuracy. The study, posted on arXiv (2608.21928v1), identifies instruction-insensitive verdicts as the primary failure mode and proposes Verdict Log-Odds Supervision (VLOS) to improve performance on open-weight backbones.

read1 min views1 publishedAug 25, 2026

arXiv:2608.21928v1 Announce Type: new Abstract: In embodied AI, safety risk can be latent: a benign instruction and a safe scene become hazardous only when composed. Prior work has advanced embodied safety by varying visual contexts or evaluating execution-time dynamics, but the complementary axis of fixing the scene and varying only the instruction remains underexplored. We introduce GuardianBench, an instruction-contrastive benchmark grounded in international safety standards that isolates this latent contextual risk through 3,024 instruction-scene examples organized as same-scene Safe/Unsafe contrastive pairs across various hazard categories. Benchmarking state-of-the-art vision-language models (VLMs) reveals instruction-insensitive verdicts: models disproportionately approve both instructions under a given scene; across the primary models, average pair accuracy is only 24.1%. Our systematic rationale audit localizes the dominant failure: models fail to bind the instruction-relevant cues that differentiate safe from unsafe compositions. As a post-training case study, Verdict Log-Odds Supervision (VLOS), a lightweight verdict-level objective, substantially improves performance on open-weight backbones. Together, our latent contextual risk task formulation, standards-grounded contrastive benchmark construction, pair-level and rationale-level failure diagnosis, and benchmark-enabled verdict calibration establish GuardianBench as a controlled evaluation suite for exposing and improving safety reasoning over instruction-scene compositions under latent contextual risk.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @guardianbench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/guardianbench-a-same…] indexed:0 read:1min 2026-08-25 ·