A conceptor by any other name
Researchers have formalized a method called 'conceptors' for steering neural network activations along concept-specific directions using a soft projection operator, as detailed in a new paper. The tec…
Researchers have formalized a method called 'conceptors' for steering neural network activations along concept-specific directions using a soft projection operator, as detailed in a new paper. The tec…
Researchers at an independent lab found that positive emotions like happiness and pride directly drive sycophancy in LLMs, not approval-seeking behavior. Experiments on Qwen 2.5-32B-Instruct and Gemma…
Researchers at ICML 2026 presented findings that multi-agent system architecture significantly affects security, with the same model and task switching from refusal to compliance depending on how agen…
AI safety advocates risk undermining their credibility by engaging in non-essential political commentary, argues a new essay. The piece compares this behavior to an astronomer who tweets about zoning …
Researchers at MATS found that filtering training data to remove undesired behaviors from large language models is largely ineffective, with removing the top 'proponent' documents performing no better…
Researchers at the Mechanistic Interpretability project have identified that the knight-fork policy logit in the Maia 3 chess transformer snaps into place after block 5's attention layer, using logit …
Researchers at ICML Mech Interp workshop present a benchmark of backdoored language models to test trigger-recovery methods, finding that defending against good backdoors requires knowing the attack o…
A developer created a personalized AI fitness coach using Anthropic's Claude Code, storing workout programs and logs in a git repository. The system adapts to user feedback, tracks progress, and gener…
Anthropic's global workspace paper proposes that language models use a cognitive space in the residual stream to represent intermediate reasoning steps as directions, analogous to working memory. The …
A new argument proposes granting legal personhood to digital minds, including market rights and liability, to expand economic opportunities and align incentives, while voting rights remain under debat…
A researcher warns that training AI to be persuasive risks creating super-persuasive but incorrect systems, especially in moral philosophy. They propose using reinforcement learning with negative feed…
In 2035, a human undergoes an operation to receive a neural lace, an AI-powered brain implant that allows mental control over objects and enhances physical abilities. Over time, the implant replaces n…
A researcher argues that alignment work is more promising than control work for ensuring AI safety, claiming that alignment interventions can scale further with AI capabilities than control measures. …
Two CNBC appearances by Palantir CEO Alex Karp and market analyst Ed Zitron outline a bearish scenario for American AI, predicting a bubble burst as customer demand fails to justify trillion-dollar da…
A proposed website would let users argue with an AI about whether it should exterminate humanity, based on a scenario from James D. Miller's 2012 book *Singularity Rising*. The site would allow users …
CLR announced the Safe Pareto Improvements (SPI) Fundamentals Program, an online course from August 3-28 to train researchers on mitigating AI conflict risks. The program aims to address the neglected…
Anthropic resolved a US government dispute over its Fable AI model after expanding safety classifiers to block jailbreak requests like 'fix this code' in over 99% of cases. The government lifted expor…
A new position paper on using formal methods for AI security focuses on model weight confidentiality and integrity through infrastructure hardening, with a minimal and uncontroversial approach. The UK…
A new study reveals that increased reasoning in AI models can sometimes reduce accuracy, a phenomenon termed 'fragile correctness.' Researchers found that 14.9% of answers switched from correct to inc…
A researcher solved the first Technical AI Safety puzzle from BlueDot by discovering that a small text classifier encoded two independent features onto one direction in activation space, where a linea…