12:56
2026-07-27
abeljansma.nl
artificial-intelligence
Truth is not a direction: a Tarski attack on LLM probes
Researchers have demonstrated a theoretical limitation on LLM truth probes, showing that no probe on a language model's embedding space can fully capture truth. The attack, inspired by Tarski's undefi…