GLM-5.2 is a win for local AI
Chinese lab Zhipu's commercial arm Z.AI released GLM-5.2, a 753-billion-parameter open-source language model with a 1-million-token context window, on June 17. The model achieves near-frontier perform…
Chinese lab Zhipu's commercial arm Z.AI released GLM-5.2, a 753-billion-parameter open-source language model with a 1-million-token context window, on June 17. The model achieves near-frontier perform…
OpenAI launched LifeSciBench on June 17, a benchmark with 750 expert-authored tasks and nearly 20,000 evaluation criteria to test AI models on real-world biological research workflows. The benchmark, …
Researchers introduced PromptMN, a pseudo-prompting domain-specific language that annotates natural language with compact typed directives to reduce context ambiguities in human-AI interactions. The l…
General-purpose large language models (LLMs) including GPT-5.2, Gemini 3.1 Pro, and Claude Opus 4.6 outperformed specialized clinical AI tools OpenEvidence and UpToDate Expert AI across three evaluati…
Researchers at arXiv propose a behavioral measure of trust between AI agents based on costly verification in a cooperative survival game. Testing six frontier model snapshots, they found that larger m…
Anthropic released Claude Fable 5, its first Mythos-class model, which achieves 80% on SWBench Pro but costs $10 per million input tokens and $50 per million output tokens. The model excels at vision …
General-purpose large language models GPT-5.2, Gemini 3.1 Pro, and Claude Opus 4.6 outperformed specialized clinical AI tools OpenEvidence and UpToDate Expert AI on medical knowledge tests, expert ali…
Google Research released Gemini-SQL2 on June 12, a text-to-SQL system built on Gemini 3.1 Pro that achieved 80.04% execution accuracy on the BIRD benchmark, the first model to surpass 80%. The system …
The Aithos Research Foundation tested 12 frontier AI models across 3,000 scenarios and found every model violated the EU AI Act and GDPR, with top performer Anthropic's Claude Opus 4.7 achieving only …
Google DeepMind researchers found that supervised fine-tuning (SFT), not reinforcement learning, drives most safety properties in Gemini models. Comparing SFT-only versions of Gemini 3.1 Pro and Gemin…
Google Research's Gemini-SQL2, built on Gemini 3.1 Pro, achieved 80.04 percent accuracy on the BIRD text-to-SQL benchmark, surpassing OpenAI and Anthropic. The technology aims to enhance natural langu…
Five AI models and one human fan were tasked with predicting every group-stage scoreline of the 2026 World Cup before kickoff. Claude Sonnet 4.6 from Anthropic scored the highest with 9 points, while …
Google Research announced Gemini-SQL2, a text-to-SQL capability powered by Gemini 3.1 Pro, which achieved 80.04% execution accuracy on the BIRD Single-Model Leaderboard, surpassing its predecessor Gem…
General-purpose frontier LLMs outperformed two specialized clinical AI tools across medical benchmarks in a Nature Medicine study published June 12. Researchers compared OpenEvidence and UpToDate Expe…
Anthropic released Claude Fable 5, calling it its most capable model and the new state-of-the-art for vision tasks, but independent benchmarks from Roboflow show the claim does not hold up on real-wor…
KiloBench launched as a new evaluation framework that measures AI coding models based on real-world production cost and performance rather than benchmark scores. The tool emerged after its creators fo…
A developer benchmarked 10 AI models across three wire formats—GCF, TOON, and JSON—and found that GCF achieved 100% comprehension and generation accuracy on frontier models like Claude Sonnet and Gemi…
Meta chief AI officer Alexandr Wang said the company's recently launched Muse Spark model is "not at the tier of the leading frontier models," calling it an "appetizer" while Meta trains stronger syst…
Researchers have developed a cost-effective method for detecting scheming behavior in AI agents by training small open-weight "deliberative monitors" that reason over a scheming specification before j…
ServiceNow released EVA-Bench Data 2.0, expanding its enterprise voice agent benchmark from one domain to three—Airline Customer Service Management, Enterprise IT Service Management, and Healthcare HR…