My AI Keeps Forgetting What We Already Decided
A developer built a git-based knowledge system called kms to solve AI coding agents' lack of persistent memory across sessions. The plugin, available for Claude Code and other agents, stores decisions…
A developer built a git-based knowledge system called kms to solve AI coding agents' lack of persistent memory across sessions. The plugin, available for Claude Code and other agents, stores decisions…
Alibaba released Qwen3.8-Max on August 2, debuting at #4 on Arena.ai's Frontend Code leaderboard, one spot above Claude Fable 5, and in a ten-task UI comparison, Qwen3.8-Max cost $3.05 versus $8.44 fo…
Moonshot AI's Kimi K3 model scored 1,679 on Arena.ai's Frontend Code leaderboard, surpassing Claude Fable 5 (1,631) and GPT-5.6 Sol (1,618) to take the #1 spot for frontend code generation. In a test …
OpenAI's GPT-5.6 Sol outperformed Anthropic's Claude Fable 5 in a planning benchmark, scoring an average of 9.10 to Fable 5's 8.04 across three backend systems, according to a test by Kilo AI. GPT-5.6…
GLM-5.2 from Z.ai shows inconsistent code review quality depending on prompt phrasing, according to a controlled test by Kilo Code CLI. On a straightforward codebase with 16 planted bugs, the model ca…
Z.ai's GLM-5.2 scored 9.0 to Moonshot AI's Kimi K2.7 Code's 8.1 in planning a feature flag service, but both built near-identical working code from the winning plan, making GLM-5.2 the preferred model…
Anthropic released Claude Fable 5, a Mythos-class model for long-running agentic work and coding. In a test comparing it to GPT-5.5 for building a feature flag service, Claude Fable 5 produced a bette…