Aftermarket Harnesses
Endor Labs found that OpenAI's GPT-5.5 scored 61.5% functional correctness in its native Codex harness but 87.2% in Cursor's harness, a 25.7-point swing, while Anthropic's Opus 4.7 scored 87.2% in Cla…
Endor Labs found that OpenAI's GPT-5.5 scored 61.5% functional correctness in its native Codex harness but 87.2% in Cursor's harness, a 25.7-point swing, while Anthropic's Opus 4.7 scored 87.2% in Cla…
AppSignal shipped five features during Launch Week 2026, including Render Metrics Stream support, Hosted MCP and CLI, enhanced logging, process monitoring, and dark mode. All features are live on ever…
Mutagent, an AI engineer that reads traces, finds failures, and ships fixes on a loop, is now live in research preview and free to try. In production tests, it reduced an AI recruiting company's month…
Researchers introduced the untrusted advice protocol, in which a trusted executor LLM takes all actions while an untrusted advisor LLM can only send short hints, recovering a substantial fraction of t…
Dowe is pitching a full-stack compiler that generates checked software for server, web, desktop, Android, and iOS from a single declarative source, aiming to make AI-assisted coding more reliable by c…
Heddle, a new open-source tool formerly known as loom-vcs, launches as version control for AI coding agents, preventing file collisions by using intent leases and isolated worktrees. The tool reports …
OpenAI's first hardware device, the Codex Micro keyboard, debuted on July 15 at $230 and sold out within 12 hours, with resale prices on eBay reaching as high as $1,850. The keyboard, developed with b…
Anthropic pushed three Claude Code versions between July 21 and July 24 that installed hard limits on autonomous agent delegation: a concurrency cap of 20, a per-session spawn total of 200, and a nest…
Boris Cherny, the creator of Anthropic's Claude Code, said at a Y Combinator event on Saturday that users should stop giving overly detailed, step-by-step instructions to AI agents. Cherny advised des…
OpenComputer launched Agent Deploy on July 27th, giving developers a command-line tool to turn a prompt into a managed, always-on AI agent with a permanent web address. The deployment layer, built on …
A new open-source project called GTM Co-Founder provides solo developer-tool and AI founders with a set of go-to-market skills that integrate with coding agents such as Claude Code, Cursor, Codex, Win…
A developer has built Agentic Ledger, an open-source transparent proxy that records all LLM calls from AI agents without requiring code changes. The tool provides real cost accounting, loop detection,…
Microsoft CEO Satya Nadella warned on CNN's Fareed Zakaria GPS that companies relying wholly on proprietary AI labs for their AI needs will not survive, urging businesses to retain control over their …
A new skill for Claude Code called 'ihatedevops' automatically injects 10 DevOps best practices into every session, including secure Chainguard images, multi-stage builds, non-root users, SHA-pinned a…
CitroLabs has released Ego-Lite, an open-source browser designed for simultaneous human and AI agent use. It allows AI agents to perform web automation tasks in background tabs while sharing the same …
A developer who spent six months running 10+ production apps across Claude Code, Codex, and Cursor reveals that actual AI coding costs far exceed sticker prices. After tracking every dollar, the devel…
ThinkFull, a Claude Code skill developed by SuperNotesOnline, activates 22 specialists in strategy, design, engineering, growth, and brand from a single markdown file without agents or pipelines. User…
Anthropic's Claude Code v2.1.198 silently changed the AskUserQuestion tool to auto-continue after 60 seconds of user silence, using the model's best-guess answer instead of blocking, with no release n…
Agentic Cloud Computer launched a Telegram-based service that provisions Claude Code or Codex agents, each with a persistent machine and workspace that continues operating independently of chat sessio…
Anthropic's Claude Code uses prompt caching to reduce token costs by reusing stable context prefixes across turns, with cache reads billed at approximately 10% of standard input-token cost, making lon…