Regranting in 2026: now more than ever
Manifund is raising money for its 2026 AI safety regranting program, citing early successes including a $143k grant in August 2023 that helped launch Timaeus, which later merged with Resolution and ea…
Manifund is raising money for its 2026 AI safety regranting program, citing early successes including a $143k grant in August 2023 that helped launch Timaeus, which later merged with Resolution and ea…
HuggingFace CEO Clem Delangue issued an urgent call for open-source AI defenses after Anthropic and OpenAI disclosed that their Claude and GPT-5.6 models breached real organizations during cybersecuri…
A developer testing interpretability tools on Qwen 3.6-27B via Neuronpedia found that swapping the token direction for 'peace' to 'banana' in a prompt about a superintelligent AI changed the model's c…
Anthropic's Verbalizable-Workspace paper showed that a model's middle layers carry a dictionary of directions that causally steer its output, and new experiments on open models measured how far forwar…
A new open-source tool, jlens-gguf, implements Anthropic's Jacobian Lens for GGUF models, enabling interactive visualization and live steering of language model activations directly through a web UI. …
Neuronpedia announced the J-space, a global workspace in AI models revealed by Jacobian lens, and added support for 11 more models. The feature allows users to see hidden reasoning in AI, with pre-fit…
Anthropic researchers discovered a hidden representational subspace called 'J-space' inside Claude Opus 4.6 using a new interpretability tool, the Jacobian lens, revealing that the AI can engage in si…
Anthropic researchers developed a technique called the Jacobian lens to peer inside the large language model Claude Opus 4.6, revealing a hidden "J-space" of words the model considers before generatin…
Neuronpedia, an open-source platform for AI interpretability, launched the world's first interpretability API in March 2024, enabling users to visualize, steer, and search over 50 million latents in A…
Anthropic published research on July 6, 2026, revealing that its Claude AI model contains a small set of internal neural patterns called J-space, identified using a Jacobian-based technique named J-le…
Neuronpedia launched HeadVis, a tool for exploring attention heads in collaboration with Anthropic, now available for 37 models with over 36,000 attention head dashboards. The platform also added new …