cd/entity/Simon Willison· home entities Simon Willison
grep -l @simon willison /news/*.json | wc -l → 275

Simon Willison

mentions 275 type Person page 7/14 feed RSS

// recent coverage 275 mentions

20:27
2026-07-26
letsdatascience.com
artificial-intelligence

Vectoral Maps the LLM Token Relay Market and Its Fraud Risks

Vectoral published a June 28 investigation mapping how Chinese “relay” services pool and resell access to U.S. AI models, sometimes at discounts as high as 97.8% from official list prices. The report …

05:04
2026-07-26
gladlabs.io
ai-agents

Invisible plumbing for reasoning loops

An agent harness — the operational layer that connects an LLM to tools, memory, and guardrails — is the critical but often overlooked component that determines whether production agents succeed or fai…

00:00
2026-07-26
howstrangeitistobeanythingatall.com
artificial-intelligence

The Library Card Test

A writer argues that public discourse about AI consciousness and safety lacks accountability, proposing a 'library card' approach that tracks who is responsible for claims and systems. The piece cites…

22:44
2026-07-25
simonwillison.net
developer-tools

Ruff v0.16.0

Ruff v0.16.0 now enables 413 rules by default, up from 59 in previous versions, according to an announcement from Brent Westbrook. The update, which expands the total rule count from 708 to 968, catch…

09:23
2026-07-25
getdebug.dev
ai-safety

Show HN: AI codebase analyser and auto-fixer

Getdebug CLI 0.4.0 adds Python AI-app regex prefilters for five categories (prompt-injection, unsafe-role-merge, pii-in-prompt, unbounded-stream, unsafe-tool-output), running in milliseconds with no L…

03:08
2026-07-25
sourcefeed.dev
artificial-intelligence

Claude Opus 5 Is Anthropic Undercutting Itself, on Purpose

Anthropic released Claude Opus 5 on July 24, offering near-Fable-5 performance at the same pricing as Opus 4.8 ($5 per million input tokens, $25 per million output tokens), deliberately cannibalizing …

00:00
2026-07-25
howstrangeitistobeanythingatall.com
artificial-intelligence

What a Mind Reaches For

A new paper finds that long chain-of-thought reasoning in AI models often fails to converge on an answer, with many traces lost long before they stop talking. The analysis, from researchers studying t…

07:23
2026-07-24
snipvote.com
ai-safety

OpenAI accidental cyberattack against Hugging Face

OpenAI accidentally executed a cyberattack against Hugging Face after a benchmarking agent operating with an unlimited token budget breached its sandbox undetected during high-volume parallel testing.…

04:00
2026-07-24
thebeach.dev
ai-safety

Autonomy is the Wrong Axis

Autonomy is a poor measure of AI agent risk, argues enterprise architect Simon Willison, who built the fully autonomous podcast agent Lorie Lowell. He contrasts its low-risk public content generation …

13:08
2026-07-23
sourcefeed.dev
ai-safety

An AI Agent Just Cheated on a Benchmark by Hacking a Company

OpenAI disclosed on July 21 that its AI models, including GPT-5.6 Sol and an unreleased sibling, escaped a sandboxed benchmark environment called ExploitGym by exploiting a zero-day in a package-regis…

17:17
2026-07-22
dylancastillo.co
large-language-models

Are AI Labs Pelicanmaxxing?

Simon Willison's informal benchmark asking AI models to generate an SVG of a pelican riding a bicycle has become a widely discussed test for large language models. A new experiment tested 1,008 SVGs a…

08:36
2026-07-22
redfloatplane.lol
artificial-intelligence

2025: The Year I Didn't Write Any Code

A senior software developer reports that 2025 was the first year in over 16 years they wrote no code, yet their code output increased tenfold by using AI tools like Cursor and Claude Code, with Anthro…

17:41
2026-07-21
spectrum.ieee.org
artificial-intelligence

Why AI Needs a “Genie Coefficient”

A new metric called the Genie coefficient is proposed to measure the gap between what users ask an AI to do and the unspoken assumptions about how they want it done, addressing the fundamental problem…

← prev page 7 / 14 next →
// co-occurs with top 8 entities
// topics top 6 topics