cd/entity/TruthfulQA· home entities TruthfulQA
grep -l @truthfulqa /news/*.json | wc -l → 13

TruthfulQA

mentions 13 type Organization feed RSS

// recent coverage 13 mentions

18:29
2026-08-12
promptcube3.com
large-language-models

Trunchbull lets you run LLM benchmarks in a browser

Trunchbull, a browser-based tool for running LLM benchmarks, lets users author tests and evaluate models in real time, with native support for the harbor authoring system and the Vercel AI SDK. It inc…

06:16
2026-08-05
pub.towardsai.net
large-language-models

Your Hallucination Benchmark Is Measuring Your Detector

A study labeling 7,440 answers from four open-weight LLMs found that more than half of the hallucination labels were incorrect, and correcting them reordered the results. The author, Priyanshi Jain, r…

04:00
2026-07-21
arxiv.org
artificial-intelligence

Diagnosing Correctness Probes under Self-Judgement Confounding

A new arXiv preprint (2607.16799v1) finds that hidden-state readouts from language models primarily encode self-judgement (SJ) rather than objective correctness (OC), with the SJ-associated direction …

08:25
2026-07-11
machinebrief.com
artificial-intelligence

Uncertainty: A Breakthrough in Neural Network Prediction

Researchers have developed a lightweight method for quantifying uncertainty in neural network predictions using two key approximations: a first-order Taylor expansion and an isotropy assumption. The a…

// co-occurs with top 8 entities
// topics top 6 topics