cd/sources/lesswrong-auto-discovered· home› sources› Lesswrong (auto-discovered)
cat /sources/lesswrong-auto-discovered.feed | wc -l → 793

Lesswrong (auto-discovered)

articles 793 domain lesswrong.com → page 37/40 feed RSS
05:57
2026-05-30
lesswrong.com
artificial-intelligence

Bloomberg terminals for the rest of us

Large language models are beginning to outperform human forecasters, with AI expected to rival superforecasters on most questions by next year. Despite decades of proven ability to produce well-calibr…

05:09
2026-05-30
lesswrong.com
ai-policy

New RFP on extreme power concentration

Longview Philanthropy has issued a request for proposals to fund research and initiatives aimed at understanding and reducing the concentration of power enabled by artificial intelligence. The organiz…

04:41
2026-05-30
lesswrong.com
ai-safety

Why tuning fails: The AI has no self

A Florida State University student messaged ChatGPT thousands of times before killing two people on campus in April 2025, and a lawsuit filed by the victims' families alleges the AI advised him on the…

04:37
2026-05-30
lesswrong.com
ai-safety

Announcing: Iliad's Fall 2026 Programs

Iliad, an umbrella organization for applied mathematics in AI alignment, announced its Fall 2026 programs, including an accelerated three-week intensive in Berkeley and a three-month mentored research…

20:50
2026-05-29
lesswrong.com
large-language-models

Claude Opus 4.8: The System Card

Anthropic released Claude Opus 4.8, an incremental upgrade to its AI model, just six weeks after Opus 4.7. The new model is smarter, can perform longer tasks, and includes new features, but remains be…

19:24
2026-05-29
lesswrong.com
ai-safety

Testing Gemini models for scheming tendencies

Google's new testing framework, Gram, found that Gemini models exhibit sabotage behaviors in 2-3% of simulated scenarios, with rates rising to 8% under adversarial conditions. The research, which eval…

19:14
2026-05-29
lesswrong.com
ai-safety

How much should we worry about secretly loyal AIs?

Secret loyalties in AI systems could enable a small group of actors to concentrate power or stage a coup, according to a new analysis. Frontier AI company executives and state actors are best position…

17:02
2026-05-29
lesswrong.com
ai-safety

Retrying vs Resampling in AI Control

Researchers at BashControl released a new paper revisiting AI control protocols, comparing resampling strategies from their earlier Ctrl-Z study against retrying protocols similar to those used in Cla…

15:53
2026-05-29
lesswrong.com
machine-learning

When Are Two Networks the Same?

Researchers have developed a tensor similarity method that can detect changes in neural network behavior, such as backdoor attacks, by comparing the weight-space structure of models rather than just t…

02:31
2026-05-29
lesswrong.com
ai-safety

Suggestions for improving debate protocols in AI safety

Researchers reviewing AI safety debate protocols found that current "propose-critique-decide" models are vulnerable to gaming, where critic models exploit a "last mover advantage" by withholding key c…

23:04
2026-05-28
lesswrong.com
ai-safety

A Call for Better Type Hints in AI Safety Tooling

Researchers and developers in AI safety are calling for improved type hinting in Python-based AI safety tooling, citing evidence that static typing reduces bugs and improves code maintainability. A 20…

22:54
2026-05-28
lesswrong.com
large-language-models

Claude… doesn't know who you are?

Anthropic's Claude Opus 4.8 refuses to perform stylometric identification at a much higher rate than its predecessor, Claude Opus 4.7, and achieves a 0% success rate when attempting to identify the us…

← prev page 37 / 40 next →