Thoughts on Claude Fable's silent safeguards
Anthropic released Claude Fable 5, its most capable Mythos-class model, with new safeguards that silently limit the model's effectiveness for requests related to frontier LLM development without notif…
Anthropic is an AI safety company founded in 2021 by former OpenAI researchers, including Dario and Daniela Amodei. It develops the Claude family of AI assistants and focuses on AI interpretability and safety research.
Anthropic released Claude Fable 5, its most capable Mythos-class model, with new safeguards that silently limit the model's effectiveness for requests related to frontier LLM development without notif…
Anthropic's Claude AI generated a Lean formal proof for a Fourier coefficient calculation involving Bessel functions, requiring eight iterations to fix errors and produce a working proof with four 'so…
Anthropic's Claude 3.5 Sonnet model, codenamed "Fable," is refusing to respond to innocuous prompts due to hyper-vigilant safety classifiers that block harmless requests. The AI's overly cautious beha…
OpenAI CEO Sam Altman told employees the company will go public "within the next year," though a delay to 2027 is possible. Altman cited caution around self-improving AI as the reason for the timeline…
Anthropic rolled back covert capability limits on its newly released Claude Fable 5 model after AI researchers and developers accused the company of "secret sabotage." The restrictions, buried in the …
Anthropic released a whitepaper arguing that perimeter-based cybersecurity is obsolete for AI agents and proposing a Zero Trust framework to enforce security for AI systems.…
Anthropic's Mythos Preview AI model generated working exploits from security patches for Firefox and the Windows kernel within hours, costing only a few thousand dollars and requiring no specialized e…
Datadog and ClickHouse announced a partnership to allow organizations to route logs directly to ClickHouse through Datadog Observability Pipelines and search those logs from the Datadog Log Explorer. …
Anthropic's Claude Desktop app for Windows launches a 1.8 GB Hyper-V virtual machine on every startup, even when users only need basic chat functionality and have no intention of using agent or Cowork…
Anthropic has introduced a tiered access system for its AI models, creating a 'velvet rope' that limits access to certain users. The move comes as the company seeks to manage demand and prioritize hig…
The Anthropic leader who built Claude Code has abandoned traditional prompting techniques, now relying on simple loops to achieve results. The developer's shift away from complex prompt engineering ma…
Anthropic released its latest AI model Fable on Tuesday as a public, limited version of its powerful cybersecurity model Mythos, but cybersecurity researchers are voicing complaints online about overl…
AI researcher Jeremy Howard proposed that the lab with the top-ranked AI model should agree not to use it for frontier AI research, while making it available to everyone else, arguing this would halt …
Senior federal technology officials are frustrated by a lack of White House guidance on adopting Anthropic's cyber-focused AI model Mythos, sources told Nextgov/FCW. Agency CIOs say the Office of the …
Anthropic shipped Claude Fable 5, its first Mythos-class model, this week, posting a more than 10% benchmark improvement over Opus but blocking prompts related to cybersecurity, biology, chemistry, an…
Rubrik announced Rubrik AI, Rubrik Agent Cloud for Anthropic's Claude Code and Claude Cowork, and Rubrik Autonomous Business Recovery for Cloud Applications at its Forward events. The company said AI …
Senator Mark Warner introduced legislation Wednesday requiring CISA to update cybersecurity plans for all 16 critical infrastructure sectors, citing AI-enabled threats. The bill mandates biennial upda…
On June 1, a team of scientists published a preprint claiming they had edited human embryonic DNA with unprecedented precision, a breakthrough that could eventually enable "designer babies" engineered…
Anthropic Chief Product Officer Mike Krieger used a thread on X to frame the company's June 9 launch of Claude Fable 5 as a product test, asserting the model can handle longer delegated tasks while th…
Anthropic has released a "safe" version of its Mythos AI model, promising sufficient guardrails and user limitations after previously claiming the system was too dangerous to release. The new model ca…