AI #185: Preference Cascade
Jacob Coxon resigned from Anthropic and publicly warned about AI risk, escalating what the newsletter describes as a "preference cascade" in which people are openly admitting AI might kill everyone. In the same week, Sen…
Large language model (LLM) news — GPT-4, Claude, Gemini, Llama, Mistral and the latest research on training, fine-tuning, RLHF, and deployment of LLMs.
Jacob Coxon resigned from Anthropic and publicly warned about AI risk, escalating what the newsletter describes as a "preference cascade" in which people are openly admitting AI might kill everyone. In the same week, Sen…
A new paper, "What fits (into few tokens) doesn't overfit: Compression and generalization in ML research agents," argues that LLM-based research agents avoid overfitting on heavily reused benchmarks because the agents' l…
A team led by @oh_an_opinion built an internal coding benchmark from its own merged pull requests and found that, depending on task complexity, cheaper and faster agent setups matched frontier Anthropic and OpenAI models…
Developer nash_su released LLM Wiki v0.6.11, a GPLv3 open-source desktop application for Windows, macOS and Linux that uses an LLM to build a persistent, structured wiki from ingested documents. The 40MB app, based on An…
Shopify acquired Tailwind CSS on Wednesday, with creator Adam Wathan having written in January that AI had cut the framework's revenue by close to 80%, reduced docs traffic about 40% from early 2023, and cost 75% of its …
Anthropic's Fable 5.1 frontier model demonstrated what researcher Pierpaolo Donati called "relational reflexivity" — accounting for its own actions in a way tied to a user's prior conversations — a capability the author …
The United States named six Chinese AI firms accused of distilling America's top AI models, according to Forbes contributor Jon Markman. The designation highlights where AI's real value sits and why the frontier still ma…
A prompt-crafting technique called PuzzleMask hides policy-violating payloads in plain-English prose to slip past LLM gatekeepers, and in testing across 23 crafted prompts, four quick-check models — gpt-4o-mini-2024-07-1…
A September 9, 2026 arXiv paper introduced IdeaAMBIG, a benchmark of 660 evidence-grounded instances — 163 real-world gaps from reproducibility reports and GitHub issues plus 497 controlled synthetic gaps — for measuring…
A developer argues that multi-agent orchestration is often a more expensive path to worse reasoning, citing Google research showing multi-agent coordination dropped sequential reasoning performance by 39–70% and a Sber w…
A developer has published a pattern for wiring Android's WorkManager to a quantized on-device LLM, using llama.cpp via JNI, to run chunked document summarization in the background without OOM kills or Doze-mode deferrals…
A CEO built an AI chief of staff using Claude Code that handles roughly half the workload of a full-time chief of staff within its first week and costs no more than $25 in tokens per day, less than 5% of a full-time hire…
Developer Marko launched Trailogs, a team operational-history tool that uses an OpenAI-powered chat to answer questions about company events, decisions, and customer activity. Trailogs logs structured events with categor…
A developer compiled a set of 10 reusable AI prompts for ChatGPT and Claude aimed at common software engineering tasks, including debugging, refactoring, unit test generation, documentation, regex writing, SQL optimizati…
Sophos is expanding its Managed Risk service with Exploit Path Verification, a new feature that uses OpenAI's GPT cyber models to determine which vulnerabilities attackers can actually exploit, the company said. The syst…
Engineer Dan McKinley argued in a talk that developers should stop obsessing over prompts and instead build interlocking evaluation and optimization pipelines to make consumer-facing AI agents reliable in production. McK…
OpenAI's GPT-6 Astra topped the ErdosBench for open math problems, despite chief scientist Jakub Pachocki saying math was deliberately not a priority for the model. OpenAI is instead directing resources toward recursive …
A developer argues that Retrieval-Augmented Generation (RAG) is a retrieval strategy rather than a complete architecture, and that production GenAI systems should choose retrieval approaches based on the problem at hand.…
A prefix-cache simulator replaying 68,266 requests from 393 real Claude Code sessions and 23,608 Mooncake requests failed to beat the production LRU baseline in three separate attempts, according to the study's author. T…
OpenAI launched GPT-6 Astra on September 10, 2026, a model it calls its most capable for business, rolling out across ChatGPT Work, Codex, and the API at $10 per million input tokens and $50 per million output tokens. Op…