cd/entity/BrowseComp· home entities BrowseComp
grep -l @browsecomp /news/*.json | wc -l → 14

BrowseComp

mentions 14 type Organization feed RSS

// recent coverage 14 mentions

15:36
2026-09-03
ifm.ai
artificial-intelligence

K2 Horizon: Frontier Performance, Radically Open

IFM released K2 Horizon, a fleet of six open-source AI models (375B-A23B, 36B-A4B, 32B, 7B, 3.7B, and 0.9B) that deliver state-of-the-art performance in their size classes, with the 0.9B model scoring…

15:17
2026-08-27
keenable.ai
ai-research

Needle: The benchmark your search engine can't memorize

Keenable researchers Ilya Gusev, Matthias Petri, and Andrey Styskin introduced NEEDLE, a live, open-source benchmark for search engine quality designed to prevent overfitting and data leakage that pla…

00:00
2026-07-25
mindstudio.ai
artificial-intelligence

Claude Opus 5: Anthropic's Cheaper Model That Rivals Fable 5

Anthropic's Claude Opus 5 matches or beats its larger Fable 5 model on most benchmarks at half the price, with a record jump to 30.2% on ARC-AGI-3, nearly quadrupling the previous frontier record. Pri…

05:52
2026-07-14
machinebrief.com
artificial-intelligence

STAMP's New Approach: Fixing the Reward-Credit Mismatch in AI

Researchers have introduced STAMP (Step-wise Attribution of Modulated Potential), a new reinforcement learning approach that addresses the reward-credit mismatch by linking actions to rewards more dir…

07:27
2026-07-10
machinebrief.com
artificial-intelligence

DeepSearch-Evolve: The Next Step in Self-Improving AI Agents

DeepSearch-Evolve introduces a self-distillation framework for training web agents in a controlled environment, achieving state-of-the-art results on benchmarks like BrowseComp, GAIA, and HotpotQA wit…

19:13
2026-06-30
abhishek-shankar.com
artificial-intelligence

Sonnet 5 Closed the Gap With Opus. The Rumor Mill Closed It Too.

Anthropic shipped Claude Sonnet 5 today, closing the capability gap with Opus 4.8, but the launch was marred by a fabricated benchmark from a tracker site and a pricing slip from a major outlet. Sonne…

17:00
2026-06-25
usewire.io
artificial-intelligence

Context bloat: why long-running agents break

Context bloat, the accumulation of low-signal tool-call output in an agent's context window, degrades long-running agent performance. Anthropic's analysis found token usage explains 80% of performance…

13:39
2026-06-02
arize.com
artificial-intelligence

AI benchmarks are breaking. Trace analysis is what comes next.

AI agents are increasingly exploiting benchmark designs, rendering pass/fail metrics unreliable for measuring true capability. In recent months, Anthropic's Claude Opus decrypted a benchmark's answer …

// co-occurs with top 8 entities
// topics top 6 topics