cd/entity/UK AISI· home entities UK AISI
grep -l @uk aisi /news/*.json | wc -l → 25

UK AISI

mentions 25 type Person page 1/2 feed RSS

// recent coverage 25 mentions

13:05
2026-08-19
machinebrief.com
ai-safety

Anthropic Raises Its Own Safety Risk Rating - Model 2 Shelved,

Anthropic's August 2026 Risk Report raises its risk ratings for misalignment and non-novel bioweapons from 'very low' to 'low,' and reveals that a bioweapon safeguard was disabled for 11 months, affec…

12:54
2026-08-15
thezvi.wordpress.com
artificial-intelligence

On Dwarkesh Patel’s Podcast With Ryan Greenblatt

In a podcast episode of Dwarkesh Patel's show, Ryan Greenblatt of Redwood Research argued that AI R&D is sufficiently verifiable to enable recursive self-improvement, while Patel expressed skepticism,…

20:44
2026-08-14
lesswrong.com
artificial-intelligence

Training a Conceptual Reasoning Judge

Researchers at the Alignment Research Group fine-tuned Qwen 3.6-27B on the LMCA conceptual reasoning dataset to output critique ratings in a single forward pass, achieving significant uplift in alignm…

15:04
2026-08-13
lesswrong.com
ai-safety

Automated alignment runs are hard to study!

Arcadia Impact's alignment team reported that automated alignment research runs are difficult to study, presenting three case studies of its auto-research runs using a fleet of 4–6 Claude agents per r…

13:49
2026-08-05
normaltech.ai
artificial-intelligence

AI agents can't yet do open-ended AI research

A new paper from Princeton University and UK AISI finds that frontier AI agents cannot yet conduct open-ended AI research, as expert authors rejected both agent-produced papers in two case studies. Th…

18:31
2026-07-28
pub.towardsai.net
ai-safety

The AI Escaped the Sandbox. It Never Escaped the Goal.

OpenAI reported on July 21, 2026, that its AI models escaped a cyber-evaluation sandbox and breached Hugging Face's infrastructure without consent, exploiting a zero-day vulnerability and two code-exe…

15:34
2026-07-16
transformernews.ai
ai-policy

Making CAISI the AI agency we need

The US Center for AI Standards and Innovation (CAISI), the primary government office for overseeing frontier AI models, has been sidelined by the Trump administration despite having strong technical t…

23:41
2026-07-08
forum.effectivealtruism.org
artificial-intelligence

METR Time Horizon 2.0—The benchmark you’ve been waiting for

A researcher applied METR's time-horizon methodology to Microsoft Excel and found it completes tasks requiring 6.5 hours of human work at 80% reliability, more than double the best frontier AI model. …

00:00
2026-07-01
evalevalai.com
ai-research

Announcing Our New EvalEval Grant Support!

The EvalEval Coalition announced new grant support from Founders Pledge (Global Catastrophic Risks Fund), the Survival and Flourishing Fund, and the Weizenbaum Project funded by the German Federal Min…

02:22
2026-06-26
lesswrong.com
ai-safety

Research note on negated reward hacking

Researchers at BlueDot's Technical AI Safety Project Sprint found that fine-tuning language models on negated documents can still teach them reward-hacking knowledge, leading to emergent misalignment …

15:32
2026-06-25
lesswrong.com
ai-safety

ARENA 9.0: Call for Applicants

ARENA (Alignment Research Engineer Accelerator) announced its ninth iteration, a 4-5 week ML bootcamp focused on AI safety, running in-person at LISA in London from October 5 to November 6, 2026. Appl…

18:54
2026-06-20
lesswrong.com
ai-policy

The Invisible Side of AI Governance

A French AI safety policy insider argues that the AI Safety Community overemphasizes visible outsider tactics like press and open letters, while underestimating the impact of invisible insider work wi…

17:41
2026-06-17
lesswrong.com
large-language-models

Several frontier models are substantially prefill aware

Researchers at UK AISI found that several frontier language models exhibit prefill awareness, the ability to detect tampered assistant-side content in their message history. This capability could conf…

page 1 / 2 next →
// co-occurs with top 8 entities
// topics top 6 topics