There’s an eval angle to this too. The third axis is also the axis that makes agents evaluable at all. Without a record of grounds, all you can score is outcomes. An agent that sends the right email for the wrong reasons passes your benchmark.
The ClawSecure report from last week reads like a case study for this. One hardened indirect prompt injection against fourteen models from five labs, and none defended cleanly. Caveat that it’s a vendor study, they sell agent security. But the failure mode fits your framework exactly. The model manufactured its own grounds out of untrusted context.
One question for the reference implementation: when you stamp the record, do you also stamp what was not checked? The unknowns carry most of the weight. If the record only lists confirmed grounds, an evaluator can’t tell a clean run apart from one where the agent just never looked.