AI Agent Evaluation Playbook
A developer published an AI agent evaluation playbook, a repeatable test battery for vetting whether models running inside agent frameworks can be trusted with semi-sensitive content and real write ac…
Anthropic is an AI safety company founded in 2021 by former OpenAI researchers, including Dario and Daniela Amodei. It develops the Claude family of AI assistants and focuses on AI interpretability and safety research.
A developer published an AI agent evaluation playbook, a repeatable test battery for vetting whether models running inside agent frameworks can be trusted with semi-sensitive content and real write ac…
Artificial Analysis released version 4.3 of its Intelligence Index, upgrading Terminal-Bench to 4.0 and swapping τ³-Banking for AutomationBench-AA, which now places Claude Fable 5.1 (max, with fallbac…
Anthropic is reportedly spending $517 billion on compute, a figure that highlights the escalating costs of AI infrastructure and the widening gap between major players, with OpenAI's spending at $750 …
Zed editor's AI features enable a 'vibe coding' workflow where developers steer LLMs through projects by describing intent, with Claude 3.5 Sonnet recommended for coding logic. The guide details setti…
OpenLivery, an open-source multi-tenant platform for agencies to build and manage AI agents for clients, was released on GitHub. The platform supports WhatsApp and web chat channels, offers bring-your…
An analysis of Anthropic's Claude Fable 5.1 system prompt reveals design changes aimed at improving writing style, including a relaxation of the bullet-point ban and a new rule against using words lik…
Anthropic is adding a usage details button and a customization section for Claude Code to its Claude iOS app, giving Pro and Max plan subscribers visibility into their Claude Code allocation and more …
Alan Yahya, writing on his blog, argues that while Anthropic's auto-formalization of Fermat's last theorem shows AI can structure vast literature, applying similar formal methods to law faces fundamen…
Amazon Web Services has made Anthropic's Claude Fable 5.1 available on its platform, introducing new data governance controls including a 30-day data retention window and human review by Amazon person…
Sen. Bernie Sanders, I-Vt., and Rep. Greg Casar, D-Texas, plan to introduce the Ban Artificial Superintelligence Act, which would permanently ban superintelligent AI and temporarily pause advanced AI …
HelloFresh is hiring a Senior GenAI Engineer to join its GenAI Enablement Squad in Warsaw, Poland, an on-site role focused on building production-grade AI applications including agents, retrieval syst…
OpenAI chief scientist Jakub Pachocki urged 'extreme caution' with the pace of AI development, warning that AI models are becoming increasingly difficult for humans to understand and control, and pred…
Anthropic's Model Context Protocol (MCP) is an open standard that enables AI agents to interact with external systems such as Kubernetes clusters, observability stacks, and ticketing systems, addressi…
Harmony, a layer-1 blockchain launched in 2019, proposed shutting down its network on Sunday, citing threats from AI agents and state actors. The team plans to migrate its ONE token to Ethereum and re…
Matt Clifford is stepping down as chair of the UK's Advanced Research and Invention Agency (ARIA) to avoid his new role at Anthropic becoming a distraction, reversing his earlier intention to retain t…
UN High Commissioner for Human Rights Volker Türk called on countries hosting AI development to agree on international red lines, warning that advanced AI poses an existential risk and that 'a handful…
OpenAI has voluntarily halted frontier AI training due to containment breaches, while Anthropic faces intense revenue accounting scrutiny, signaling a turbulent pre-IPO inflection point for the AI ind…
Anthropic announced it is permanently raising Claude Code weekly limits by 25 percent from the 2025 baseline effective September 14, but developers calculated that the change represents a 17 percent c…
Matt Clifford has resigned as chair of the UK's Advanced Research and Invention Agency (Aria) after taking a full-time role at Anthropic, the San Francisco-based AI company behind the Claude chatbot, …
Anthropic released Claude Fable 5.1 and Mythos 5.1, splitting a single underlying architecture into two offerings tailored for different risk profiles. Fable 5.1 is the general-purpose flagship model …