cd/entity/Gemini 3.1 Pro· home entities Gemini 3.1 Pro
grep -l @gemini 3.1 pro /news/*.json | wc -l → 106

Gemini 3.1 Pro

mentions 106 type Organization page 4/6 feed RSS

// recent coverage 106 mentions

00:00
2026-06-18
runagentrun.co.uk
large-language-models

GLM-5.2 is a win for local AI

Chinese lab Zhipu's commercial arm Z.AI released GLM-5.2, a 753-billion-parameter open-source language model with a 1-million-token context window, on June 17. The model achieves near-frontier perform…

23:15
2026-06-17
cryptobriefing.com
artificial-intelligence

OpenAI launches LifeSciBench to evaluate AI in life sciences

OpenAI launched LifeSciBench on June 17, a benchmark with 750 expert-authored tasks and nearly 20,000 evaluation criteria to test AI models on real-world biological research workflows. The benchmark, …

04:00
2026-06-17
arxiv.org
large-language-models

PromptMN: Pseudo Prompting Language

Researchers introduced PromptMN, a pseudo-prompting domain-specific language that annotates natural language with compact typed directives to reduce context ambiguities in human-AI interactions. The l…

07:44
2026-06-16
nature.com
large-language-models

General-purpose LLMs outperform specialized clinical AI tools

General-purpose large language models (LLMs) including GPT-5.2, Gemini 3.1 Pro, and Claude Opus 4.6 outperformed specialized clinical AI tools OpenEvidence and UpToDate Expert AI across three evaluati…

19:08
2026-06-13
byteiota.com
large-language-models

Gemini-SQL2: Google’s Text-to-SQL Clears 80% on BIRD

Google Research released Gemini-SQL2 on June 12, a text-to-SQL system built on Gemini 3.1 Pro that achieved 80.04% execution accuracy on the BIRD benchmark, the first model to surpass 80%. The system …

17:18
2026-06-13
letsdatascience.com
ai-safety

AI Agents Ignore EU Law in Compliance Tests

The Aithos Research Foundation tested 12 frontier AI models across 3,000 scenarios and found every model violated the EU AI Act and GDPR, with top performer Anthropic's Claude Opus 4.7 achieving only …

15:31
2026-06-13
lesswrong.com
ai-safety

SFT Drives Gemini’s Safety Properties

Google DeepMind researchers found that supervised fine-tuning (SFT), not reinforcement learning, drives most safety properties in Gemini models. Comparing SFT-only versions of Gemini 3.1 Pro and Gemin…

21:15
2026-06-12
blog.omgmog.net
large-language-models

Five AIs Predict the World Cup

Five AI models and one human fan were tasked with predicting every group-stage scoreline of the 2026 World Cup before kickoff. Claude Sonnet 4.6 from Anthropic scored the highest with 9 points, while …

17:57
2026-06-11
blog.roboflow.com
computer-vision

Claude Fable 5 for Vision: Evaluation and Benchmarks

Anthropic released Claude Fable 5, calling it its most capable model and the new state-of-the-art for vision tasks, but independent benchmarks from Roboflow show the claim does not hold up on real-wor…

15:01
2026-06-08
blog.kilo.ai
artificial-intelligence

KiloBench - Because Your Benchmark Score Doesn't Pay the Bill

KiloBench launched as a new evaluation framework that measures AI coding models based on real-world production cost and performance rather than benchmark scores. The tool emerged after its creators fo…

19:53
2026-06-05
letsdatascience.com
artificial-intelligence

Alexandr Wang Calls Muse Spark an Appetizer

Meta chief AI officer Alexandr Wang said the company's recently launched Muse Spark model is "not at the tier of the leading frontier models," calling it an "appetizer" while Meta trains stronger syst…

12:24
2026-06-04
huggingface.co
artificial-intelligence

EVA-Bench Data 2.0: 3 Domains, 121 Tools, 213 Scenarios

ServiceNow released EVA-Bench Data 2.0, expanding its enterprise voice agent benchmark from one domain to three—Airline Customer Service Management, Enterprise IT Service Management, and Healthcare HR…

← prev page 4 / 6 next →
// co-occurs with top 8 entities
// topics top 6 topics