cd/entity/TransformerLens· home entities TransformerLens
grep -l @transformerlens /news/*.json | wc -l → 6

TransformerLens

mentions 6 type Organization feed RSS

// recent coverage 6 mentions

05:32
2026-08-10
lesswrong.com
ai-safety

How to be an AI safety research engineer

A practical guide for aspiring AI safety research engineers advises focusing on specific issues, proactive networking, and a year-long upskilling process, with Python, PyTorch, and linear algebra as e…

16:03
2026-07-21
lesswrong.com
ai-safety

Steering Blackmail Through a Model's "Emotional State"

A case study on Gemma 3 12B reveals that the model's decision to blackmail an executive is not linearly decodable until late in its reasoning, peaking at layer 19 with 0.74 AUROC, and that steering an…

23:04
2026-05-28
lesswrong.com
ai-safety

A Call for Better Type Hints in AI Safety Tooling

Researchers and developers in AI safety are calling for improved type hinting in Python-based AI safety tooling, citing evidence that static typing reduces bugs and improves code maintainability. A 20…

// co-occurs with top 8 entities
// topics top 6 topics