unwind ai- Posts
- Microsoft's new skill tunes your agents in Claude Code, Codex, and Cursor
+ Free tiers of 28 LLM providers in one API #
Start here ↓
Your coding agent can now tune and optimize your other agents. Microsoft's Agent Lightning v1.0.1 installs as a skill in Claude Code, Codex, GitHub Copilot, or any other agent harness.
Give it an editable agent and a benchmark, and it reworks prompts, tools, workflows, model choice, and reasoning settings against scored results instead of vibes.
Every change is measured, so the eval you keep meaning to write is the price of entry, and writing it is how you finally learn whether last Tuesday's prompt edit helped.
🚀** Shipped** #
Xiaomi announced the AI Cube, a desktop machine specified at 1.2 TB/s memory bandwidth. Local serving hits a memory-bandwidth wall long before a FLOPs wall, so that is the number worth watching on a box like this. No first-party page resolved, so every spec is the thread's claim.ithome.com
Headless Tools launched SaaS tools with no user interface at all, built to be called by agents directly. Everyone else is bolting an agent API onto a product designed for a human; this one starts from the agent. hdls.tools**Grok 4.6 is now available inside Hermes at 50% off through the Nous Research portal **for the next week. You can also use an existing SuperGrok or X Premium+ subscription. If you use Hermes for longer agent runs, this is a cheap week to test whether Grok earns a slot in your routing table.x.com/SpaceXAI
Dactyl builds native mobile apps from a description, runs them live in a browser simulator, and can ship to TestFlight without a Mac or Xcode. Comes with live simulator: Apple sign-in, camera, sensors, and payments can work while the app is still being generated. dactyl.dev
🧠** Worth Knowing** #
OpenAI cut GPT-5.6 Luna pricing by 80% and Terra pricing by 20% across the API, ChatGPT Work, and Codex. Luna is now $0.20 / $1.20 per million input/output tokens, and Terra is $2 / $12, so the cheap end of the GPT-5.6 family is now much cheaper for agent loops. openai.com
Claude Team and admins can now authorize MCP connectors once for the whole org through their identity provider. This fixes the annoying rollout problem where every user had to do their own OAuth flow before Claude could use the same tools. For companies, MCP just got easier to deploy and audit.support.claude.com
Your agent does not have to wait for the model to finish before starting the slow work. Speculative Programmatic Tool Calling predicts likely calls from partially written code, then launches search, sub-agents, or APIs early so they run alongside token generation. Worth testing if tool latency is the bottleneck in your harness.alexzhang13.github.io
**Artificial Analysis launched benchmarks for phone-sized models across intelligence and real mobile-device inference. **Great for comparing small models on tool use, reasoning, speed, latency, and memory instead of relying on vendor claims or “works on my phone” demos.artificialanalysis.ai
**Local AI image generation in Microsoft Paint is not fully local. **Paint and Photos still send prompts to a remote server for moderation, then hide a server-issued GUID inside the generated image. If your local pipeline promises nothing leaves the device, check this.xusheng.dev
**Tool sandboxing is not enough if the model server is exposed. **Boyd Kane walks through how an LLM could reach its host machine through the inference engine itself, not shell access or normal tool calls. Worth reading if you run local or self-hosted inference.boydkane.com
🔧** Clone and Run** #
**Clone & Run of the Day ****freellmapi puts the free tiers of 28 LLM providers behind one OpenAI-compatible **/v1
** endpoint**. Point an existing client at it and you can test more models without making budget the blocker. Keep it for experiments, not user-facing production traffic.github.com/tashfeenahmed/freellmapi
claude-obsidian turns an Obsidian vault into a self-organizing Claude knowledge graph. You drop in sources, Claude reads and links them, and the output stays as plain Markdown files you own. Nice if you like Karpathy’s LLM Wiki idea but want it inside a vault you already use.github.com/AgriciDaniel/claude-obsidian
Codebuff’s freebuff splits terminal coding work across four specialized agents, and each role can use a different model. That is the practical version of multi-agent coding: spend the expensive model where it matters, use cheaper ones where it does not. If one generalist agent keeps thrashing in your terminal, this is worth trying.github.com/CodebuffAI/freebuff
effective-html is a set of agent skills for making better HTML artifacts: wireframes, prototypes, plans, and diagrams. Most agents can produce something that renders, but “renders” is not the same as “looks good and explains the idea.” Skills are a clean way to fix that once and reuse it across sessions.github.com/plannotator/effective-html
OCR It is a Chrome extension that pins a screen region and OCRs paginated documents locally with bundled Tesseract. Basically, it is for the annoying PDF or document viewer that will not let you copy text. Pin the area, hit the hotkey, and send the extracted text to your LLM.github.com/thiagotigaz/ocr-it
This GPU checker tells you whether a specific GPU can run a specific LLM, then estimates fit, speed, and stack placement. That turns the most common local-LLM argument into a lookup instead of a forum rabbit hole. The interface is Korean-first, but the calculator is still worth bookmarking.jaeseok614.github.io
**x64dbg-mcp-server exposes x64dbg controls over HTTP to any MCP client. **Breakpoints, stepping, memory reads, and register dumps become things an assistant can drive directly instead of asking you to paste debugger output. Small repo, big idea if you do reverse engineering or low-level debugging.github.com/duty1g/x64dbg-mcp-server
Awesome LLM Apps** (134k+ 🌟 ) is a curated collection of 100+ AI Agents, Agent skills, and RAG apps.** It covers models from OpenAI, Anthropic, Google, and open-source models like GLM, DeepSeek, and Qwen that you can run locally on your computer. (Now accepting GitHub sponsorships)
📊** By the Number** #
Number of the Day****Peter Walker from OpenRouter reports that GPT-5.6 Sol now makes up 50%+ of US business spend inside OpenAI’s model family on OpenRouter. OpenAI’s frontier model is already taking most of the spend, while Luna and Terra just got cheaper for the work that does not need Sol. The routing question is now pretty direct: what actually needs Sol, and what can move to the cheaper models?x.com/PeterJ_Walker
Amazon raised hardware prices by 60%, blaming the memory shortage, in the same week Nvidia told customers to expect AI server increases above 15%. If you have been comparing rented inference against owning the box, that comparison moved twice in seven days.techcrunch.com
**Hugging Face is fielding acquisition interest that would value it at roughly $13 billion. **The default hosting layer for open weights being priced for sale is a supply-chain question for anyone whose pipeline starts with a from_pretrained call. Exploratory interest, not an agreed deal, and no acquirer is named.businessinsider.com
That’s all for today. Come back tomorrow for the next batch of AI tools, model drops, agent repos, and weird benchmarks worth your time.
If you found one thing to try, share the issue with someone who ships.
The best prompt engineers aren't typing. They're talking.
Power users figured this out early: speaking a prompt gives you 10x more context in half the time. You include the edge cases, the examples, the tone you want — because talking is fast enough that you don't skip them.
Wispr Flow captures everything you say and turns it into clean, structured text for any AI tool. Speak messy. Get polished input. Paste into ChatGPT, Claude, Cursor, or wherever you work.
89% of messages sent with zero edits. 4x faster than typing. Works system-wide on Mac, Windows, and iPhone.