LaughBench
LaughBench, a new benchmark created by Taylor G., tests AI models' ability to generate novel, funny jokes as a measure of general intelligence. Frontier models GPT-5.6 Sol and Fable occasionally made …
LaughBench, a new benchmark created by Taylor G., tests AI models' ability to generate novel, funny jokes as a measure of general intelligence. Frontier models GPT-5.6 Sol and Fable occasionally made …
A user who previously preferred GPT over Claude criticizes Anthropic's Opus 5 model as 'awful' and a 'slop-cannon,' claiming it fails to deliver on explicit goals, makes bad decisions and large mistak…
A developer comparing C++ and Rust for AI/ML code with LLMs found that while Rust's Burn framework is competitive on GPU, it performs poorly on CPU compared to C++ with GGML. The experiments, publishe…
Cybersecurity researchers report that safety guardrails from OpenAI and Anthropic are preventing legitimate offensive security work to find unknown vulnerabilities and develop exploit tools, according…
AI-bedrijven zoals OpenAI en Anthropic gebruiken volgens onderzoekers angstaanjagende doemscenario's om de kracht van hun technologie te benadrukken, terwijl ze die uiteindelijk toch op de markt breng…
Moonshot AI's open-source Kimi K3 model, after a week of real project use, matches or surpasses Claude Opus in instruction following and long coding sessions, and rivals Fable in data analysis, but su…
Moonshot AI's release of the 2.8 trillion-parameter Kimi K3 model has triggered White House accusations of model distillation from Anthropic's Fable model, with Treasury Secretary Scott Bessent threat…
Cursor launched Cursor Router, an intelligent model router for Teams and Enterprise customers that selects an AI model before each coding request, targeting frontier-level performance while reducing s…
Anthropic mathematician Levent Alpöge announced a disproof of the 87-year-old Jacobian conjecture in a tweet, crediting the AI model Fable for its help. Separately, a researcher used ChatGPT prompts s…
Scientists have developed a 'Ghost Font' that uses motion to write messages readable by humans but indecipherable by AI, leaving even advanced models like Fable and GPT Sol 5.6 Ultra baffled. Develope…
A Hacker News user seeking a writing model free of 'claudeisms' and sloppiness reports that Opus 5 and Fable have disappointed, with Fable being 'ridiculously more expensive' and only slightly better.…
A pseudonymous hacker known as Pliny the Liberator claims to have broken every major AI model at once, including GPT-5.6 Sol Ultra, Opus 5, and Fable, with a universal jailbreak announced July 24, tho…
A security researcher known as Pliny the Liberator claims to have discovered a universal jailbreak technique effective on all AI models, including heavily guardrailed flagships like Opus 5, GPT-5.6 So…
Anthropic has released Opus 5, a new AI model priced at half the cost of its Fable sibling, and the company says it does not require data retention.…
Anthropic rolled out Opus 5, a new version of its model focused on token efficiency rather than a major capability leap, offering performance slightly ahead of Opus 4.8 and OpenAI's GPT-5.6-Sol on ben…
Nearly four years after ChatGPT's release, large language models still suffer from the same core flaws — poor math, outdated knowledge, short memory, and toxicity — that were present in GPT-3.5, accor…
The White House accused Chinese AI company Moonshot of distilling Anthropic's Fable model to build its Kimi K3, and of obtaining restricted Nvidia servers through Thailand to evade US export controls.…
The White House accused Moonshot AI of using a sophisticated distillation attack to exploit Anthropic's Fable model for its Kimi K3, with Treasury Secretary Scott Bessent warning that sanctions and En…
Moonshot AI is accused of distilling Anthropic's Fable model to train its own model, raising questions about the use of synthetic data and intellectual property in AI development. The dispute highligh…
AI guardrails designed to prevent malicious use are hindering offensive cybersecurity researchers, who say the restrictions block legitimate work like finding and confirming software vulnerabilities. …