Introducing: interp-engine 🚀🔎
Decode Research has open-sourced interp-engine, a high-performance interpretability engine built from scratch to run production Neuronpedia workloads, including Jacobian Lens, NLAs, circuit tracing, a…
Decode Research has open-sourced interp-engine, a high-performance interpretability engine built from scratch to run production Neuronpedia workloads, including Jacobian Lens, NLAs, circuit tracing, a…
Chinese labs released 10 open-weight AI models in the past 30 days, including DeepSeek-V2.5 (236B parameters, Apache 2.0) and Qwen-1.5-110B-Chat, challenging US counterparts like Llama 3 and Gemma 2 w…
A developer reports that Mistral NeMo 12B is the best local LLM for a Mac Mini M4 Pro with 24GB of unified memory, balancing reasoning and tool calling, but warns that adding a reranker like qllama/bg…
Google has released Gemma 2, the next generation of its open models, featuring a 27B parameter version that offers performance competitive with models more than twice its size while being efficient en…
Google released Gemma 2, an open model family with 9B and 27B parameter sizes, featuring architectural changes like hybrid attention and Grouped-Query Attention for improved inference efficiency. The …
New research spanning 11 models including Qwen 2.5, Gemma 2, and Llama 3.2 reveals that larger language models systematically shift their evaluation-awareness to earlier network layers, indicating the…
Google's Gemma 2 models demonstrate that architectural efficiency can deliver competitive performance with fewer parameters. The 27B model rivals models twice its size through hybrid attention, Groupe…
Open-weight AI models are now 3–6 months behind frontier cloud models in benchmark performance, but the gap is closing fast enough that local AI has become a viable infrastructure decision for cost, p…