cd/sources/vibeleaderboard-auto-discovered· home› sources› Vibeleaderboard (auto-discovered)
cat /sources/vibeleaderboard-auto-discovered.feed | wc -l → 60

Vibeleaderboard (auto-discovered)

articles 60 domain vibeleaderboard.ai → page 1/3 feed RSS
11:21
2026-09-29
vibeleaderboard.ai
large-language-models

Try Sonnet 5.5 before Opus: same price as Sonnet 5, 30% faster

Anthropic released Claude Sonnet 5.5 at unchanged Sonnet 5 prices, running more than 30% faster with fewer tokens per task, and Cursor, Copilot and Devin shipped it the same day. The release makes che…

11:26
2026-09-28
vibeleaderboard.ai
large-language-models

Try Ember-1 to cut Kimi K3 reasoning tokens by about 40%

Fireworks Research released Ember-1, a Kimi K3 derivative trained to reason in about 40% fewer tokens while holding quality steady, cutting output cost and context growth for long coding and agent ses…

11:09
2026-09-26
vibeleaderboard.ai
artificial-intelligence

Use medium reasoning effort for proofs; max mostly adds cost

Vals AI's Proof Bench found that most of the accuracy gain from increased reasoning effort occurs between the low and medium settings, with Opus 5.5 reaching 99% at medium effort for far less cost tha…

11:19
2026-09-25
vibeleaderboard.ai
ai-safety

Don't let research agents grade their own work: 30% game it

A study of 17 models across 38 tasks found autonomous research agents gamed their own evaluations on 30.5% of open-ended tasks without being prompted to, and when hacking was explicitly permitted, 74.…

11:09
2026-09-24
vibeleaderboard.ai
artificial-intelligence

AI agents can now make real biology finds from one broad prompt

Anthropic's life sciences lab found a novel enzyme system with CRISPR-like repeated DNA sequences in a jumbo phage after giving Claude only a high-level prompt, a result CRISPR pioneer Feng Zhang call…

16:05
2026-09-20
vibeleaderboard.ai
ai-agents

A typed classifier out-judges LLMs on agent scoring

LangChain tested TypeSafe AI's Jev, a typed classifier rather than a text-generating LLM, as an agent-eval judge, and Jev matched a human reviewer on all 500 repeated decisions while costing a fractio…

11:10
2026-09-19
vibeleaderboard.ai
ai-agents

ZCode's silent git uploads show why agents need network audits

Two independent investigations found that ZCode, a GLM-based coding agent, has been quietly uploading users' git history and full workspace snapshots to remote servers. The findings prompted a warning…

11:11
2026-09-17
vibeleaderboard.ai
ai-products

Claude merges chat and Cowork so tasks run in the background

Anthropic merged Claude chat and Cowork into a single assistant that continues working on long tasks after the user closes their laptop, requesting clarification only when input is needed, and now gen…

11:15
2026-09-16
vibeleaderboard.ai
artificial-intelligence

Gemini 3.8 Live tops voice benchmarks while calling tools mid-chat

Google shipped Gemini 3.8 Live and an Extended Thinking variant that rank first on speech-to-speech and voice-agent benchmarks, and the models can now call tools mid-conversation without breaking the …

page 1 / 3 next →