Does Your LLM Trust You?
A study by an anonymous researcher, conducted as part of Neel Nanda's MATS 10.0 stream, found that linear 'trust' vectors extracted from the residual streams of Llama-3.2-3B-Instruct and Llama-3.1-8B-…
A study by an anonymous researcher, conducted as part of Neel Nanda's MATS 10.0 stream, found that linear 'trust' vectors extracted from the residual streams of Llama-3.2-3B-Instruct and Llama-3.1-8B-…
A new study from arXiv finds that linear readouts trained on the final-token hidden state of large language models can decode whether diagnostic evidence supports, challenges, or fails to address a ca…
A study presented at the 2026 Annual Meeting of the Special Interest Group on Discourse and Dialogue (SIGDIAL) finds that large language models used as role-playing agents exhibit a trade-off between …
A new study from researchers at Forethought Foundation finds that training language models to be risk-averse on low-stakes gambles (prizes up to $100) can generalize to astronomically high stakes (pri…
A new benchmark evaluates KV-cache optimization techniques—quantization, pruning, and merging—for long-context LLM serving, finding that compression ratio alone poorly predicts end-to-end performance.…
Researchers propose SPSD, an edge-based pipeline that compresses user prompts using a small language model before sending them to a cloud LLM, reducing input tokens by an average of 99.9 per call whil…