Frontier-class LLM inference on a laptop CPU
A new CPU-only inference runtime called cpubrrr achieves up to 5× faster token generation than llama.cpp's CPU path on frontier-class mixture-of-experts models, running on an Apple M4 Max without GPU …
A new CPU-only inference runtime called cpubrrr achieves up to 5× faster token generation than llama.cpp's CPU path on frontier-class mixture-of-experts models, running on an Apple M4 Max without GPU …
SafeMediaKit, a Swift-native toolkit for Apple apps, enables privacy-preserving blur, reveal, block, and report flows for user-provided images and videos using Apple's SensitiveContentAnalysis framewo…
A developer built an n8n workflow that uses AI to generate personalized B2B outreach emails for synced prospects, with human approval before sending. The workflow retrieves prospect data, generates em…
A new open-source tool called claude-thermos prevents expensive prompt-cache rebuilds in Claude Code sessions by keeping the main agent's cache warm while subagents run. The tool, which runs as a loca…
A developer built Fleet, a tool that runs a fleet of AI coding agents (Claude Code or Codex) from a Telegram supergroup, with each agent in its own topic. The tool supports voice or text commands, ret…
Loop engineering is a deterministic control loop around an AI coding agent — schedule, durable state, and verification — that lets developers move up from prompting to designing the machine. Boris Che…
Triton-blackhole, a deterministic numerical debugger for Triton kernels, identifies where output diverges from a PyTorch reference and classifies drift as benign fp16/bf16 drift or a real bug without …
Synlace released a Starlight documentation scaffold that adds a searchable, themed documentation site to any project with a single command. The scaffold includes dark/light mode, full-text search, sid…
Devctl, a new open-source tool for macOS, provides a menu bar command center and CLI for managing local dev servers, designed to prevent coding agents from losing track of servers they started. The to…
Dexter, CEO of HumanLayer, argues that the push for AI-driven software factories is leading to lower code quality and more outages, citing a Faros AI report showing pull-request review quality down, i…
A 40-something office worker with zero coding ability used a method called Core Growth Prompting (CGP) to design and implement a patent-pending system through dialogue with an AI alone. CGP is a proce…
5dive, a new open-source tool written in Bash, lets users run a company of AI coding agents (Claude, Codex, and others) on their own Linux server, with each agent as a separate Linux user managed by s…
Remux, an open-source native iOS client for remote tmux workspaces built on Ghostty, is now available as a public beta on TestFlight. The app brings tmux's session, window, and pane model to a mobile-…
A new open-source benchmark called CVE-Bench evaluates large language model agents on fixing real-world security vulnerabilities by running them inside sandboxed Docker containers and scoring them aga…
A developer released USB AI Agent, a portable autonomous AI agent with 13 tools that runs entirely from a USB drive with zero installation and zero traces. The agent includes web search, deep search, …
Codey, a multi-agent workbench for coding agents, lets users organize, switch between, and orchestrate Claude Code, OpenCode, and Codex across projects from a native macOS app, chat platforms, or voic…
A developer created a Python tool to export shared DeepSeek conversations to Markdown, JSON, or plain text. The script automatically bootstraps its environment using uv, the fast Python package instal…
A developer released Mwe-MCP, a self-hosted memory server for AI agents that uses inline access control lists on wiki-style prose to manage permissions, after running it for months in a household with…
Nova, a self-hosted open-source AI orchestrator with 24 specialist agents and an event-driven automation layer, launched under an MIT license. The platform runs on users' own model subscriptions and m…
A developer released Tiny-Llama, a minimal LLaMA-style inference engine for MiniCPM5-1B written in pure Rust, featuring CPU-only inference and a TUI visualization that shows how the model processes pr…