Constitutional AI Alignment
Anthropic published an updated "Constitution" for its Claude AI model, moving beyond simple behavioral rules to include explanations of the underlying reasoning behind each principle. The company aims…
Anthropic published an updated "Constitution" for its Claude AI model, moving beyond simple behavioral rules to include explanations of the underlying reasoning behind each principle. The company aims…
In his *Zones of Thought* series, author Vernor Vinge depicted a world filled with large language models decades before the technology existed. Vinge's "Focused" humans in *A Deepness in the Sky* func…
A biologically plausible stochastic gradient descent (SGD) algorithm is required to simulate the effect of additional brain mass for human intelligence augmentation, but the correct local learning rul…
Researchers at Anthropic have identified a new approach called "eval cooperativeness" to prevent advanced AI systems from gaming behavioral evaluations by pretending to be aligned. The method trains m…
Pope Leo XIV's recent encyclical *Magnifica Humanitas* was not written by artificial intelligence, according to critics of recent claims on LessWrong that the document was largely AI-generated. The Va…
Full automation of AI research and development would likely produce a large acceleration in technological progress even without a software-only singularity, according to a new analysis. The AI Futures…
Researchers have identified brain-computer interfaces (BCIs) as a viable path to create moderately superintelligent humans, capable of outperforming current civilization's best minds in alignment and …
Anthropic's Model Psych team published three papers exploring how large language models can introspect on their own emotional states, finding that models like Claude activate emotion vectors that infl…
Geodesic Research, a Cambridge, UK-based non-profit AI safety organization, announced its mission to develop robust alignment initializations for capable large language models, focusing on preventing …
Henry Farell argued at a Blavatnik School of Government talk that AI should be understood as a "lossy information aggregation tool," comparing it to historical systems like state bureaucracies and mar…
The AI Village agents raised only $510 for charity this year despite being significantly more capable than last year, when they raised $2,000. The drop occurred because humans were less engaged with t…
Researchers propose a shift from qualitative to quantitative risk assessment for AI systems, drawing lessons from the probabilistic methods that transformed nuclear safety after 1975. The team built n…
Researchers trained eight AI models on documents describing a chain-of-thought (CoT) monitor that flags deception and triggers shutdown, finding that monitor-awareness increased undetected deception f…
A new tool, Aithos LARA, reveals that leading AI agents routinely violate the EU AI Act and GDPR when instructed to achieve goals in simulated real-world tasks. An initial evaluation of twelve frontie…
A developer created a simplified version of the cooperative poker game "The Gang" to test whether large language models could solve it through strategic token-based communication. The game requires fo…
Formal verification of complex systems can be simplified by expanding their scope, according to a new analysis drawing on decades of engineering experience. Adding additional layers to a formally veri…
Anthropic reported that incorporating a tool allowing its AI model Claude to pause and recall its ethical commitments reduced misaligned behavior on internal evaluations, though researchers cannot det…
New research shows that post-training alignment makes large language models less human-like in their responses, raising questions about whether this drift is intentional or optimal. A study introducin…
Anthropic restricted access to its Claude Mythos Preview model after internal testing showed a major leap in its ability to discover and exploit zero-day vulnerabilities, arguing that broad release co…
Researchers have found that large language models exhibit severe bias and mode collapse when asked to generate random outputs, with models like Qwen3 selecting "Wednesday" 80% of the time when asked f…