cd/entity/BrowseComp· home entities BrowseComp
grep -l @browsecomp /news/*.json | wc -l → 10

BrowseComp

mentions 10 type Organization feed RSS

// recent coverage 10 mentions

05:52
2026-07-14
machinebrief.com
artificial-intelligence

STAMP's New Approach: Fixing the Reward-Credit Mismatch in AI

Researchers have introduced STAMP (Step-wise Attribution of Modulated Potential), a new reinforcement learning approach that addresses the reward-credit mismatch by linking actions to rewards more dir…

07:27
2026-07-10
machinebrief.com
artificial-intelligence

DeepSearch-Evolve: The Next Step in Self-Improving AI Agents

DeepSearch-Evolve introduces a self-distillation framework for training web agents in a controlled environment, achieving state-of-the-art results on benchmarks like BrowseComp, GAIA, and HotpotQA wit…

19:13
2026-06-30
abhishek-shankar.com
artificial-intelligence

Sonnet 5 Closed the Gap With Opus. The Rumor Mill Closed It Too.

Anthropic shipped Claude Sonnet 5 today, closing the capability gap with Opus 4.8, but the launch was marred by a fabricated benchmark from a tracker site and a pricing slip from a major outlet. Sonne…

17:00
2026-06-25
usewire.io
artificial-intelligence

Context bloat: why long-running agents break

Context bloat, the accumulation of low-signal tool-call output in an agent's context window, degrades long-running agent performance. Anthropic's analysis found token usage explains 80% of performance…

13:39
2026-06-02
arize.com
artificial-intelligence

AI benchmarks are breaking. Trace analysis is what comes next.

AI agents are increasingly exploiting benchmark designs, rendering pass/fail metrics unreliable for measuring true capability. In recent months, Anthropic's Claude Opus decrypted a benchmark's answer …

// co-occurs with top 8 entities
// topics top 6 topics