Models Are Getting Dumber on Purpose
Reasoning models are deliberately trading world knowledge for reasoning skill, with GLM-5.2 scoring 99.2% on AIME 2026 using about 40 billion active parameters per token, while factual recall remains …
Reasoning models are deliberately trading world knowledge for reasoning skill, with GLM-5.2 scoring 99.2% on AIME 2026 using about 40 billion active parameters per token, while factual recall remains …
A new AI model called continuous-query limited memory language model (CO-LMLM) combines minimal cost with high factual precision by storing knowledge in an external knowledge base and fetching it as n…
Google's AI Overviews, used by over two billion monthly users, are accurate about 90% of the time, according to a New York Times analysis by startup Oumi. The 10% error rate translates to tens of mill…
An analysis by Oumi of Google's AI Overviews found that while accuracy improved from 85% on Gemini 2 to 91% on Gemini 3 on the SimpleQA benchmark, the rate of ungrounded claims among correct answers i…
Researchers introduced Decoupled Search Grounding (DSG), a vendor-agnostic architecture that separates search from reasoning in LLM agents, enabling independent control over retrieval policy, provider…
A new study from researchers testing adapter composition in large language models found that geometry-aware merging strategies, which enforce orthogonality or directional independence in parameter upd…