More than just code review
Simon Willison argues that the key skill for using coding agents is confidently instructing them and verifying changes, noting that reviewing every line of code is not the most effective validation method.
Full-text search across 57993 articles. Combine with topic and date filters; results sorted by relevance.
Simon Willison argues that the key skill for using coding agents is confidently instructing them and verifying changes, noting that reviewing every line of code is not the most effective validation method.
Alibaba's Qwen 3.8, a 2.4 trillion-parameter sparse Mixture-of-Experts model with 95 billion active parameters, has closed the general reasoning gap with US frontier labs, scoring 92.6 on GPQA Diamond and 86.6 on Termina…
Daniel Vaughn released Huzzah, an experimental editor that lets developers write pseudocode and synchronizes it to real source code on save, with the pseudocode persisted as a record of intent. Vaughn, who has worked wit…
A developer pitted OpenAI's Codex CLI against Anthropic's Claude Code in a Liar's Dice tournament, with Claude Code winning all three best-of-3 series. The engineer built a tamper-proof MCP-based setup to ensure fair pla…
Anthony "chovy" Ettinger is live streaming a mosh coding session at 1:45am Pacific on pairux.com, working with his agentic coding CLI, moshcode, on real repositories and failing tests. The session is accessible via the d…
Claude Code, Anthropic's coding assistant, improved an enterprise AI agent's performance on 500 complex queries by automatically analyzing failures and rewriting prompts in a continuous loop, achieving in 48 hours what w…
A solo developer used Claude Code, an agentic coding tool, to automatically produce a narrated product demo video for ClinTrialFinder, a free clinical-trial matching tool for cancer patients. The agent drove a live web a…
Anthropic published a technical guide on August 15, 2026, explaining how users can reduce costs and extend session limits in Claude Code, its terminal-based coding assistant, by leveraging prompt caching. The guide detai…
Zhipu AI released GLM-5.3 on August 14, 2026, through its GLM Coding Plan, claiming it is the strongest open-weights coding model, with a 50% improvement over GLM-5.2 achieved through post-training alone. The company say…
Zoom Workplace users should update to version 7.1.5 or 7.0.6 to patch three annotation flaws that could let a meeting participant execute code on another participant's client, according to Zoom security bulletins publish…
Lovable, a vibe-coding platform that lets users build full-stack web apps from plain English descriptions, has raised $400 million in Series C funding at a $13.3 billion valuation, led by Menlo Ventures and co-led by EQT…
Rust is the largest beneficiary among programming languages of coding agents, according to a Google engineer's analysis. The language's strict compiler acts as an automatic verifier that complements AI coding, while its …
A developer built a benchmark to test whether AI coding models change their behavior when given a safety skill, measuring the Keelwright Score (KDS) as execution rate times discrimination rate. Results show poolside/lagu…
A developer warns that AI coding assistants can leave secrets in git history, as they may commit API keys or other credentials to make code run. The developer recommends scanning full git history before pushing to produc…
Claude Code, an AI coding agent from Anthropic, solved a complex math problem that had stumped a developer for an extended period by identifying a symmetry in variables that simplified a multi-step iterative process into…
Plannotator, a free, MIT/Apache-2.0 licensed local code review tool, launches with a diff viewer for git, jj, and Perforce, supporting local diffs, commits, branches, worktrees, and GitHub/GitLab PRs. The tool allows man…
RayTally's analysis of Claude Code's cross-session messaging highlights a coordination gap in parallel coding agents: developers become human message buses, manually transferring context and resolving conflicts. The prop…
Linejudge, an independent verification harness for coding agents, found that 8 out of 8 agent-authored patches passed machine checks but only 6 out of 8 actually fixed the target issues in sqlite-utils, with two patches …
Lovable has become the first AI coding agent platform to earn AIUC-1 certification, the industry's first security, safety, and reliability standard for AI agents. Developed with input from Stanford, MIT, MITRE, and the C…
Perso AI released an open-source FFmpeg plugin that lets coding agents subtitle, translate, dub, and clip videos from natural-language commands, supporting files, folders, and YouTube/TikTok URLs. The plugin, available f…