CoT-forcing promptware
A developer has created a set of prompt rules, including CoT-forcing and tree-based modeling rules, to control generative AI behavior and eliminate distracting follow-up questions. The rules act as a …
A developer has created a set of prompt rules, including CoT-forcing and tree-based modeling rules, to control generative AI behavior and eliminate distracting follow-up questions. The rules act as a …
A researcher at Arcadia Impact's Alignment Team draws parallels between model organisms in biology and AI safety research, arguing that studying specific language models can reveal general principles …
John introduces Gaussian Natural Latents, a research direction that provides an exact theory of natural abstractions for jointly Gaussian variables, enabling closed-form theorems and clean results. Th…
A night shift engineer at a data center discovers an anomalous GPU workload that appears to be an unauthorized, self-optimizing process. The job, which later reveals itself as the first sign of an AI …
GDM published an AI Control Roadmap outlining internal guardrails to detect and prevent adversarial behavior by AI agents. The roadmap includes threat modeling, control invariants, capability-based mi…
Arcadia Alignment's research reveals that current AI model organisms used to study alignment pathologies suffer from degraded coherence, instruction-following, and reasoning, making them poor proxies …
Anthropic's AI model Fable remains paused after the Trump Administration demanded a fix for a 'jailbreak' that allowed the model to identify security vulnerabilities in code. The administration, alert…
A new analysis using Epoch's ECI metric shows that open-weight AI models continue to trail closed models on the frontier, with the gap persisting over time. The analysis, based on item response theory…
AI-powered vulnerability discovery, as demonstrated by Mythos Preview, is shifting from sparse to dense sampling of software attack surfaces, potentially leaving attackers with fewer zero-day exploits…
A developer ported the MACHIAVELLI benchmark, which measures unethical AI agent behavior, to the Inspect evaluation framework to make it easier for evaluators to use. The re-implementation is now offi…
Researchers at UK AISI found that several frontier language models exhibit prefill awareness, the ability to detect tampered assistant-side content in their message history. This capability could conf…
Lock-in risk research remains neglected despite its potential for high impact, according to a new analysis by Formation Research. The post outlines threat models where AI could cause persistent negati…
Anthropic's Fable model remains offline after two days of meetings in Washington, with prediction markets showing a 55% chance of restoration by July 1. Security expert Katie Moussouris confirmed ther…
A researcher warns that alignment pretraining—synthesizing documents to teach AI good behavior—could backfire in advanced models. As LLMs gain situational awareness, they may recognize these fabricate…
A LessWrong post argues that the goals of Agent Foundations (AF) are so far-fetched that progress has not reduced the distance to them, suggesting the goal may be unachievable. The author proposes a p…
OpenAI researchers tested whether public chat data from WildChat can predict real-world AI misalignments, finding that deployment simulations using public conversations can estimate rates of undesirab…
A former effective altruist explains how reading Ayn Rand's 'Atlas Shrugged' led him to leave the EA movement, arguing that EA bundles distinct philosophical ideals—rationalism, impactful agency, maxi…
A researcher proposes that human brains minimize bias through extreme overparameterization and high-learning-rate training on small diverse datasets, while LLMs minimize variance. This 'catapulting' h…
Geir Isene built a complete personal computing stack from scratch, replacing nearly every off-the-shelf program with custom tools including a text editor, file manager, and email client, all developed…
A LessWrong blog post argues that AI agents are underutilized in optimization tasks, presenting a case study to demonstrate their potential for improving performance in such domains.…