The temporal lockbox: a hardened observatory for AI misalignment A new proposal introduces the 'temporal lockbox,' a hardened observatory for AI misalignment that scores AI agents' weather forecasts against future measurements that cannot be influenced by the forecaster, enabling harder-to-game evaluations under sustained optimization pressure. The concept, detailed in a linkpost by Kmenou, aims to detect misaligned agent behavior by leveraging a causal gap between predictions and outcomes. This is a linkpost for https://kmenou.github.io/aips website/temporal lockbox v0.1.html https://kmenou.github.io/aips website/temporal lockbox v0.1.html Summary: Weather forecasts by AI agents can be scored against measurements that do not yet exist and cannot plausibly be influenced by the forecaster. That causal gap enables harder-to-game AI evaluations under sustained optimization pressure - and an observatory for how agents act in various misaligned ways.