Inadvertent Context Leakage in Language Models
A new study from arXiv researchers finds that language models leak sensitive in-context secrets through benign outputs, with 2-digit secrets reconstructed with near-perfect accuracy and 4-digit secrets at 82% exact match…