cd/entity/BBH· home› entities› BBH
grep -l @bbh /news/*.json | wc -l → 4

BBH

mentions 4 type Organization feed RSS

// recent coverage 4 mentions

21:39
2026-09-01
huggingface.co
artificial-intelligence

BenchMIRT: What are LLM benchmarks actually measuring?

The Allen Institute for AI introduced BenchMIRT, a new method for auditing LLM benchmarks at the level of individual prompts, which uses multidimensional Item Response Theory to separate the underlyin…

04:00
2026-05-27
arxiv.org
artificial-intelligence

SPEAR: Code-Augmented Agentic Prompt Optimization

Researchers introduced SPEAR (Sandboxed Prompt Engineer with Active Roll-back), a code-augmented agentic optimizer that autonomously rewrites prompts using a Python sandbox for structural error analys…

// co-occurs with top 8 entities
// topics top 6 topics