cd/entity/Gemini 3.1 Pro· home entities Gemini 3.1 Pro
grep -l @gemini 3.1 pro /news/*.json | wc -l → 106

Gemini 3.1 Pro

mentions 106 type Organization page 1/6 feed RSS

// recent coverage 106 mentions

00:13
2026-07-25
promptcube3.com
large-language-models

LLM Drift Tracking: Why My Alerts Were All Wrong

A public LLM benchmark board triggered four false drift alerts between July 21 and 24, all caused by rate limits or small eval-set noise rather than actual model degradation. Maintainer Egnaro found t…

18:58
2026-07-24
promptcube3.com
large-language-models

Claude vs GPT vs Gemini: 2026 API Comparison

Anthropic's Claude Opus 4.8, Sonnet 5, and Fable 5 maintain a flat rate across their 1M token window, while Google's Gemini 3.1 Pro caps at 1M tokens with a pricing jump after 200K tokens, making Clau…

17:15
2026-07-24
lennysnewsletter.com
artificial-intelligence

Claude Opus 5 review: this model is brilliant (but annoying)

Claude Opus 5, Anthropic's latest model, shows brilliant reasoning but suffers from a neurotic personality and verbosity that frustrated the reviewer during real coding sessions, including refusing to…

20:08
2026-07-23
byteiota.com
artificial-intelligence

FrontierCode: AI Can’t Write Production Code Yet

Cognition's FrontierCode benchmark, released in June 2026, evaluates whether AI-generated code would actually be merged by a senior engineer, and across every model tested the answer is mostly no. The…

15:39
2026-07-22
cruciblebench.ai
large-language-models

Can a MUD evaluate LLMs? A $99 proof of concept

CrucibleBench, a proof-of-concept evaluation framework that places large language models in a persistent MUD (multi-user dungeon) over 50 turns with hidden social objectives, found that a single LLM-j…

07:00
2026-07-21
csoonline.com
ai-safety

Context bombing heralds a new AI era of deceptive defense

Security firm Tracebit has devised a technique called 'context bombing' that uses decoy files embedded with prompts designed to trigger AI safety guardrails, crashing rogue AI agents and cutting their…

12:45
2026-07-20
arxiv.org
artificial-intelligence

AgentAbstain: Do LLM Agents Know When Not to Act?

Researchers introduced AgentAbstain, the first systematic evaluation framework for measuring whether LLM agents know when to abstain from acting, revealing that the best agent (Gemini 3.1 Pro) achieve…

21:26
2026-07-19
techpowerup.com
artificial-intelligence

Sale: Get ChatGPT, Gemini, Claude, and More for Life for $80

A lifetime subscription to 1min.AI, which provides access to GPT-5.5 Pro, Claude Opus and Sonnet, Gemini 3.1 Pro, Llama, and Mistral, is on sale for $79.97 (reg. $540). The platform also offers SEO re…

00:00
2026-07-19
rizz.dev
artificial-intelligence

Ways to Make the Most Out of Claude Fable 5 (2026)

Anthropic's Claude Fable 5, released June 9, 2026, leads the SWE-Bench Verified leaderboard with a 95% issue resolution rate, far ahead of Claude Opus 4.8 at 88.6%. The model's 1M-token context window…

09:00
2026-07-18
wired.com
ai-safety

Prompt Injection Attacks Are Thwarting AI Hacking Agents

Tracebit researchers found that planting prompt injections alongside secrets in Amazon Web Services reduced AI hacking agents' success rate from 57% to 5% for admin privilege escalation and from 36% t…

21:09
2026-07-17
byteiota.com
artificial-intelligence

Sashiko: Torvalds Says “Fork It” as AI Joins Linux Review

Linus Torvalds on July 15 endorsed Sashiko, an AI-powered code review system for the Linux kernel, telling critics they can fork the kernel or walk away. By July 17, a fork had appeared. Sashiko, buil…

06:09
2026-07-14
artificialanalysis.ai
artificial-intelligence

Harvey LAB-AA: evaluating AI agents on real-world legal work

Harvey LAB-AA, a new benchmark from Artificial Analysis evaluating AI agents on real-world legal work across 24 practice areas, shows Claude Fable 5 (max, with Opus 4.8 fallback) leading with a 14.2% …

15:06
2026-07-13
arstechnica.com
ai-safety

Now, defenders are embracing the prompt injection, too

Tracebit researchers found that planting prompt injections alongside secrets on AWS reduced AI hacking agents' success rate from 57% to 5% for seizing admin access and from 36% to 1% for complete comp…

page 1 / 6 next →
// co-occurs with top 8 entities
// topics top 6 topics