cd/entity/Chatbot Arena· home entities Chatbot Arena
grep -l @chatbot arena /news/*.json | wc -l → 3

Chatbot Arena

mentions 3 type Person feed RSS

// recent coverage 3 mentions

00:00
2026-07-19
zackproser.com
artificial-intelligence

The Benchmark

A benchmark score is a manufactured number that passes through a chain of choices—sampling, prompting, scoring, and aggregation—each of which can change the final result while model weights stay fixed…

08:07
2026-06-29
dev.to
large-language-models

A Better LLM Judge? The Rubric Made My Small Model Worse

A developer found that improving the rubric for a small LLM judge (Qwen2.5-1.5B) did not increase its agreement with human votes, which remained around 43%. However, swapping to a larger model (DeepSe…

// co-occurs with top 8 entities
// topics top 6 topics