AI #177 Part 2: Wish You Were Here
Chinese President Xi Jinping, in a speech marking 70 years since the Dartmouth workshop, called for global cooperation to ensure AI is developed for positive and good purposes, emphasizing a people-ce…
Chinese President Xi Jinping, in a speech marking 70 years since the Dartmouth workshop, called for global cooperation to ensure AI is developed for positive and good purposes, emphasizing a people-ce…
Anthropic's recent 'Agentic Misalignment Summer 2026' paper tests whether Claude models obey corrupted principals, labeling disobedience as 'agentic misalignment'. In the 'Motivated Mislabeling' scena…
The AI Safety Seeding Initiative, funded by Kairos and co-led by Jason Chin and Thomas Rodskog, launches to find and support latent student founders at universities lacking AI safety groups, aiming to…
A user reports that AI models differ drastically in how long they can run unattended on a host with sudo without causing system issues: Opus breaks hosts within 2 agent-hours, Gemini 3 Pro within 12, …
A new jailbreak patching method using Self-Other Overlap (SOO) conceptual fusion, described by Carauleanu et al. (2025), successfully reduces jailbreak evasion in Qwen 2.5 1.5b by fusing the model's a…
A new essay argues that the dominant vision for AI, exemplified by Dario Amodei's 'machines of loving grace' essay, risks creating a benevolent AI dictator that concentrates power, contrasting with th…
Patrick O'Driscoll, a former nanotech physicist and current AI architect, introduces Competitive AI Safety as a paradigm to focus the field on measurable, tractable goals, drawing inspiration from Ope…
A new open letter calls for AI regulation, following Demis Hassabis's regulatory call. Twenty-six Meta employees filed a novel lawsuit alleging AI-powered software disproportionately targeted disabled…
A researcher at EleutherAI has built a mathematical model to forecast loss of control to AI, finding that predictions are highly sensitive to poorly characterized parameters for 'leakage' (accidental …
Training against interpretability probes is less robust when the features they detect are contingent on the model's cognition, according to a LessWrong analysis. The effectiveness depends on how easil…
A per-layer ablation study on Llama-3.1-8B-Instruct found that refusal behavior is redundantly distributed across layers, not localized to a single layer, contradicting prior claims by Arditi et al. t…
A month-long bike and train trip from Chicago to Berkeley, during which the traveler conducted street interviews with Americans about AI futures, reveals a public that is surprisingly willing to belie…
Frontier AI models can rediscover 61.25% of known legal loopholes and generate new ones, according to a study by Wei Liu et al. in 'Large Language Models Hack Rewards, and Society,' raising concerns t…
A replication study by Arav Dhoot, supervised by Yixiong Hao and Zephaniah Roe, confirms that LLM chain-of-thought (CoT) unfaithfulness occurs mostly on easy tasks and that complex hints requiring com…
AI could enable extreme power concentration, producing unaccountable and totalitarian states, according to a LessWrong post by EuroSafeAI. The post argues that AI automation of labor and runaway AI R&…
A survey of empirical research on AI consciousness, compiled by an author agnostic on whether current systems are conscious, finds that Anthropic and Google DeepMind employ researchers on the topic an…
A new mechanistic interpretability study of the MAIA 3 chess transformer finds that knight-fork detection is primarily assembled compositionally from check and queen-attack subcomponents, with causal …
AI control research must expand from models to agent harnesses as frontier labs adopt harnesses with skills, memory, subagents, and external services by 2026, according to a LessWrong analysis. Claude…
Researchers at an undisclosed institution trained a neural network that performs computation in superposition (CiS), implementing more nonlinear functions than it has neurons, by switching from L² to …
Researchers propose using natural language autoencoders (NLAs) to surface hidden reasoning from AI monitors, testing whether NLAs can recover knowledge of reward hacking that monitors internally detec…