Explore With an Agent, Replay Without One
Deltix, an agentic UX testing tool for mobile apps, debuted on Hacker News with a bumpy launch but introduced a design pattern that the agentic-testing space is converging on: AI at authoring time, de…
Deltix, an agentic UX testing tool for mobile apps, debuted on Hacker News with a bumpy launch but introduced a design pattern that the agentic-testing space is converging on: AI at authoring time, de…
A bug in Sentry's JavaScript SDK double-counts tool calls in streaming mode for Google's Gemini integration, recording each invocation twice with inconsistent parameter keys (`args` vs. `arguments`). …
Llmfit, a Rust CLI and TUI created by Alex Jones, the developer behind k8sgpt, right-sizes local LLMs to a user's hardware by detecting RAM, CPU, and GPU and scoring hundreds of models on quality, spe…
A solo developer's weekend project NexusOS, built almost entirely with Cursor, failed not on code bugs but on deployment issues such as environment topology, migration ordering, and provider gaps, acc…
A team from the University of Cambridge's Machine Learning Systems Lab, working with NVIDIA, Flower Labs, MBZUAI, and Inria, introduced the Red Queen Gödel Machine, a co-evolutionary system that impro…
Anthropic has begun watermarking all text outputs from new Claude models globally with an invisible statistical watermark, meeting the EU AI Act's August 2 deadline, with no opt-out and no regional ge…
A 2026 arXiv audit of real-world Java workflows found only 4% compliance with least-privilege permission controls, and model-generated GitHub Actions workflows often inherit over-permissive habits fro…
A Cambridge-led team with collaborators at NVIDIA, Flower Labs, MBZUAI, and Inria introduced the Red Queen Gödel Machine (RQGM), a framework that co-evolves an AI agent and its evaluator to prevent se…
Google shipped Gemini 3.7 Flash on August 13, three weeks after 3.6 Flash, with benchmark gains of 43.6% vs 34.4% on FrontierCode 1.1 Main, 65.3% vs 49.0% on DeepSWE v1.1, and 30.4% vs 17.0% on Automa…
Anthropic's published system-prompt history for Claude models shows a sawtooth pattern of growth and pruning, not a monotonic ratchet, with the Opus 5 prompt at about 3,200 words after shrinking for t…
Linux 7.2, released by Linus Torvalds on August 16, introduces cache-aware load balancing for chiplet-based servers like AMD EPYC and Intel Xeon 6, but the feature is opt-in via CONFIG_SCHED_CACHE. Th…
A dev.to write-up describes a setup where Claude Code is routed through a LiteLLM proxy to cheaper DeepSeek models via OpenRouter, a pattern that trades away data governance and supply-chain security.…
Frontier AI models are activating far fewer parameters per token than three years ago — Zhipu's GLM-5.2 activates about 40 billion, Alibaba's Qwen3.5 runs 17 billion active, and DeepSeek V4-Flash gets…
Ivan Gavran, a researcher on the Quint specification language, argued in a blog post that AI coding agents have shifted the economics of formal verification, making the 1979 critique by De Millo, Lipt…
Stripe has finalized a deal to buy OpenRouter for more than $7 billion, about five times the $1.3 billion valuation OpenRouter raised at in May, according to Bloomberg. The acquisition, which Stripe h…
GLM-5.2 scores 99.2% on AIME 2026 with 40 billion active parameters, while Qwen3.5's 9B model hallucinates on 82% of knowledge questions in Artificial Analysis's Omniscience benchmark, reflecting a de…
Dmitry Grinberg's critique of RISC-V, published on Hacker News, argues the ISA's design is flawed but will still dominate cheap microcontrollers, while Armstrong Subero, an embedded engineer from Trin…
Wild Static, a Show HN project, hosts a single AI named 'Static' with a shared, never-reset memory that all visitors can influence, inverting the per-user memory silos used by ChatGPT and Claude. With…
Langfuse 4.14.4 and OpenTelemetry can be used to trace a two-agent Claude pipeline, capturing every LLM call, tool invocation, token count, and dollar cost in a single nested trace timeline. The tutor…
Researchers at the Max Planck Institute for Intelligent Systems, ELLIS Tübingen, and ETH Zürich built LittleLearner, a 5B-parameter language model pretrained exclusively on 88 billion tokens of U.S. e…