Origin Part 19: The Number Was Wrong
A developer discovered that an 88% eval score on a brain layer for Origin was invalid due to data leakage: 23 of 26 held-out probes were also in training data. After fixing the overlap, the true score…
A developer discovered that an 88% eval score on a brain layer for Origin was invalid due to data leakage: 23 of 26 held-out probes were also in training data. After fixing the overlap, the true score…
A developer has been running an autonomous agent on a 16GB M1 Mac for three months, using small models in the 9B/E4B class by choice rather than economic necessity. The developer argues that small mod…
A developer built PrivateScribe, an offline AI note-taking app that runs entirely in the browser using WebGPU, with zero data leaving the device. The tool handles note summarization, email drafting, a…
A developer building an LLM observability service discovered five silent ways cost tracking can be wrong, including missing token counts for streaming requests and miscalculated prompt caching costs a…
A developer built QAMap, a local CLI tool that reads commits and diffs to describe behavior changes and highlight missing QA questions around failures, boundaries, and state changes. The tool runs loc…
Developer Maneshwar explores fusuma, a Markdown-to-slides tool that converts a single flat file into presentations, PDFs, and social preview cards. The tool uses HTML comments for slide metadata, supp…
Researchers from Zhejiang University and Alibaba presented a study at ICML 2026 showing that logically inconsistent prompts can cause reasoning models to enter long, unproductive internal loops, effec…
A developer has released the Nervous System, an MCP server that enforces mechanical rules on autonomous LLM agents to prevent common failures like losing context or taking irreversible actions. The pu…
A developer argues that the biggest problem with AI coding assistants is not the AI itself but how code is written. By adding metadata like file paths, usage locations, and feature READMEs, developers…
A developer ran 1,790 measurements across five AI engines and found that most AI visibility dashboards report misleading numbers. Branded prompts inflate scores, and different engines produce conflict…
A developer presents a 7-point framework for evaluating AI engineers in 2026, arguing that traditional signals like Kaggle experience or PyTorch knowledge no longer predict who can ship reliable AI sy…
A developer relies on two free browser-based image tools that require no installation: an online image compressor that processes files locally to avoid uploads, and ClarifyPix, which runs AI in-browse…
A developer tracked time spent correcting AI-generated code over two weeks and found that 78% of correction time was due to project-specific mismatches that could have been prevented by rules, not gen…
A developer outlines a framework for building effective human-in-the-loop systems for AI coding agents, arguing that approval prompts should be reserved for actions a human can realistically catch, wh…
A developer describes using git worktrees to run parallel Claude Code agents without file conflicts. Each agent gets its own worktree and branch, isolating file changes and preventing mid-run overwrit…
A developer building a no-code platform that generates code with AI explains that generating code is easy, but the hard part is making it fit into the app and remain maintainable over time. The platfo…
AgentGateway, the Rust-based AI proxy from the AI Agent Infrastructure Foundation, now supports RFC 8693 OAuth 2.0 Token Exchange to securely delegate user identity to AI agents. The feature replaces …
A developer is building an AI platform for CAD/CAE engineering that goes beyond traditional LLMs by combining AI agents, CAD feature understanding, engineering knowledge graphs, and simulation-aware d…
A developer has released doerkit, an open-source assessment-first course that uses an LLM to grade written answers against a rubric, wrapped in spaced cumulative review. The system, built by Michael T…
A developer built Simmerz, an iOS app, and detailed the creation of the Rork AI App Optimizer, an open-source tool that automates static analysis and runtime heuristics to refactor resource-heavy code…