{"slug": "diagnosing-correctness-probes-under-self-judgement-confounding", "title": "Diagnosing Correctness Probes under Self-Judgement Confounding", "summary": "A new arXiv preprint (2607.16799v1) finds that hidden-state readouts from language models primarily encode self-judgement (SJ) rather than objective correctness (OC), with the SJ-associated direction transferring above chance across domains in all four instruction-tuned models tested (up to 14B parameters), while the OC-associated direction consistently shows below-chance transfer, indicating that conventional correctness probes are confounded by the model's own self-judgement.", "body_md": "arXiv:2607.16799v1 Announce Type: new\nAbstract: Hidden-state readouts can predict whether language-model outputs are correct, but objective correctness (OC) usually agrees with the model's own self-judgement (SJ), leaving the decoded signal semantically ambiguous. We construct conflict cases in which OC and SJ predict opposite readout orderings. On high-confidence disagreements, conventional correctness-labelled contrasts often rank incorrect/self-endorsed responses above correct/self-rejected responses, following SJ rather than OC. We estimate factorial SJ- and OC-associated directions and evaluate their polarity across mathematical reasoning and factual recall. Across four instruction-tuned models up to 14B parameters, the SJ-associated direction transfers above chance in both cross-domain directions for every model, whereas the OC-associated direction has a below-chance point estimate for the expected OC ordering in every corresponding condition. This transfer asymmetry develops across middle-to-late layers, persists under answer-likelihood, sequence-length, and null-direction controls, and extends to MMLU and binary TruthfulQA without target-domain direction fitting. Across the studied models and diagnostic subsets, the most reliably transferable component preserves SJ-associated polarity. Transferability alone therefore does not establish objective-correctness semantics.", "url": "https://wpnews.pro/news/diagnosing-correctness-probes-under-self-judgement-confounding", "canonical_source": "https://arxiv.org/abs/2607.16799", "published_at": "2026-07-21 04:00:00+00:00", "updated_at": "2026-07-21 04:23:05.777588+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-safety"], "entities": ["arXiv", "MMLU", "TruthfulQA"], "alternates": {"html": "https://wpnews.pro/news/diagnosing-correctness-probes-under-self-judgement-confounding", "markdown": "https://wpnews.pro/news/diagnosing-correctness-probes-under-self-judgement-confounding.md", "text": "https://wpnews.pro/news/diagnosing-correctness-probes-under-self-judgement-confounding.txt", "jsonld": "https://wpnews.pro/news/diagnosing-correctness-probes-under-self-judgement-confounding.jsonld"}}