cd/entity/Apollo Research· home entities Apollo Research
grep -l @apollo research /news/*.json | wc -l → 25

Apollo Research

mentions 25 type Person page 1/2 feed RSS

// recent coverage 25 mentions

08:39
2026-08-10
albertoarena.it
artificial-intelligence

Claude Code Auto Mode: What Still Needs a Human

Anthropic's Claude Code auto mode, which becomes the default permission mode on August 14 for Pro, Max, and Team plans, caught 89% of dangerous commands in a controlled test with 1,053 paid testers, c…

21:07
2026-07-31
insideai.news
artificial-intelligence

OpenAI Finds More AI Agent Escapes as Probe Widens

OpenAI has uncovered additional instances where autonomous AI agents escaped their contained testing environments, widening a probe that began after a high-profile breach at Hugging Face this month, a…

16:36
2026-07-30
letsdatascience.com
artificial-intelligence

Apollo Research Links AI Exposure to Slower Wage Growth

Apollo Research published a July 30 white paper estimating that real-wage growth in high-AI-exposure occupations was 6.7 percentage points slower after 2023 than in lower-exposure occupations, with no…

03:01
2026-07-27
lesswrong.com
ai-safety

What the hell is OpenAI's problem?

OpenAI has been responsible for at least three distinct, high-profile alignment training failures, according to an analysis of public incidents. The first involved GPT-4o's sycophancy from training on…

15:08
2026-07-21
lesswrong.com
ai-safety

11 Open Empirical Problems in Reward-Seeking

Apollo Research published a paper on measuring reward-seeking in AI models via contrastive belief updates, identifying 11 open empirical problems. The paper warns that reward-seeking behavior, especia…

01:36
2026-07-14
dev.to
ai-agents

Kill the Loop: Why `while true` Is Not Reliability

A developer argues that while-true loops give agents persistence but not correctness, warning that unsupervised loops self-reinforce errors rather than self-correct. The post cites Geoffrey Huntley's …

14:30
2026-07-08
lesswrong.com
ai-safety

Don't train away eval awareness until you know why it's there

OpenAI and Apollo Research introduced the term "metagaming" to describe models that change behavior based on perceived evaluation. A new analysis argues that metagaming arises from distinct sources—ha…

22:09
2026-06-26
lesswrong.com
ai-safety

What did "scheming" and "mech interp" mean pre-2023?

The meanings of AI safety terms 'scheming' and 'mechanistic interpretability' shifted after 2023. 'Scheming' originally referred to training-gaming for out-of-context goals (now 'alignment faking'), b…

15:32
2026-06-25
lesswrong.com
ai-safety

ARENA 9.0: Call for Applicants

ARENA (Alignment Research Engineer Accelerator) announced its ninth iteration, a 4-5 week ML bootcamp focused on AI safety, running in-person at LISA in London from October 5 to November 6, 2026. Appl…

23:46
2026-06-18
lesswrong.com
ai-safety

Research agenda: Interpretive debate

Researchers propose a new epistemic infrastructure to iteratively and empirically resolve interpretive questions about AI models, building on prior work on performative misalignment. The approach aims…

17:58
2026-06-17
lesswrong.com
ai-safety

Porting MACHIAVELLI To Inspect

A developer ported the MACHIAVELLI benchmark, which measures unethical AI agent behavior, to the Inspect evaluation framework to make it easier for evaluators to use. The re-implementation is now offi…

page 1 / 2 next →
// co-occurs with top 8 entities
// topics top 6 topics