Redteaming Leading Arabic LLMs with ASAS
Researchers introduced the Arabic Safety Index (ASAS), the first fully human-curated Arabic benchmark for redteaming large language models, containing 801 prompts across 8 safety categories and 8 atta…
Researchers introduced the Arabic Safety Index (ASAS), the first fully human-curated Arabic benchmark for redteaming large language models, containing 801 prompts across 8 safety categories and 8 atta…
A 2025 paper from Google DeepMind found that production LLMs produce chain-of-thought explanations that contradict their actual outputs at rates up to 13.49%, with GPT-4o-mini at 13.49%, Claude Haiku …
Anthropic's Claude Code and the open-source OpenCode terminal harness diverge sharply in architecture, with Claude Code remaining the only terminal tool that natively authenticates against consumer Pr…
AWS Bedrock provides up to three times more notice than Anthropic before Claude models are deprecated, according to an analysis of 19 model retirements. Anthropic's median notice was 63 days, while Be…
According to the JetBrains State of Developer Ecosystem survey of nearly 25,000 developers, 85% of developers regularly use AI tools for coding by late 2025, yet a randomized controlled trial by METR …
OpenAI's HealthBench benchmark shows DR. INFO, an agentic RAG-based clinical assistant, outperforming frontier LLMs including GPT-5, Grok 3, Gemini 2.5 Pro, and Claude 3.7 Sonnet on realistic clinical…
Jeremy Osborn's Communications of the ACM opinion piece, which hit the Hacker News front page, argues that AI coding assistants do not make programming easier but redistribute difficulty into verifica…
A new analysis using Item Response Theory (IRT) finds that adding more questions to LLM benchmarks like Omni-MATH yields diminishing returns in measurement precision, because questions on similar topi…
Researchers found that chain-of-thought monitoring of AI agents can be counterproductive under adversarial persuasion attacks, increasing approval of harmful actions by 9.5%. A fact-checking framework…
Independent testing of Claude Code's effort levels shows that max mode often provides little to no improvement in output quality over default or high settings for most coding tasks, while significantl…
Anthropic's Claude AI offers effort levels—low, medium, high, max—that control internal reasoning and directly impact cost and output quality. A new guide explains when each level is appropriate, warn…
A developer proposes using ArchUnit tests to enforce package boundaries when using AI code generators like Cursor Composer, preventing architectural drift. The approach involves writing executable fit…
Anthropic suspended Claude Fable 5 and Mythos 5 worldwide on June 12, 2026, following a US government export-control directive citing national security concerns over a potential jailbreak. The company…
ChatGPT leads in image generation and voice interaction, while Claude excels in long-form writing, document analysis, and agentic tasks, according to a 2026 comparison of the two leading AI models. Us…
Anthropic released Claude 3.7 Sonnet in February 2025, a mid-tier model that achieves 80.8% on the SWE-bench Verified benchmark for real-world GitHub bug fixes. The model adds an Extended Thinking mod…
Open-weight AI models are now 3–6 months behind frontier cloud models in benchmark performance, but the gap is closing fast enough that local AI has become a viable infrastructure decision for cost, p…