Can a MUD evaluate LLMs? A $99 proof of concept
CrucibleBench, a proof-of-concept evaluation framework that places large language models in a persistent MUD (multi-user dungeon) over 50 turns with hidden social objectives, found that a single LLM-j…
CrucibleBench, a proof-of-concept evaluation framework that places large language models in a persistent MUD (multi-user dungeon) over 50 turns with hidden social objectives, found that a single LLM-j…
A developer's analysis of coding agent costs reveals a 63× price spread across models, from $19/month for Qwen3.5-Flash to $1,200/month for GPT-5.6 Sol, based on a fixed workload of 90M input and 25M …
Mistral AI offers the Le Chat assistant and a model family including Mistral Large 3, Mistral Small 4, and Mistral Medium 3.5, with flagship models featuring a 256K token context window and 675B total…
Mistral launched early access to its next frontier Mixture-of-Experts model, which CEO Arthur Mensch described as "fat but sparse" to close the quality gap with OpenAI and Anthropic while offering Apa…
Sierra Research introduces the Context Engine and meta-distillation to address specification failure in enterprise agents, improving retrieval and action-check pass rates on benchmarks like τ³-bench, …
Mistral AI teased a new sparse 'fat' model family for early access in July, but a viral meme about a 'Le Chaton Fat' model with 24T-30T parameters was debunked as an internet hoax. The company's actua…