cd /news/ai-tools/hutch-local-code-reviews-in-emacs-fo… · home › topics › ai-tools › article
[ARTICLE · art-145854] src=kitallis.in ↗ pub= topic=ai-tools verified=true sentiment=· neutral

hutch: local code reviews in emacs for the mildly disenfranchised

Developer adjaecent released Hutch (magit-hutch), a local code-review interface for Emacs that runs LLM-powered reviews on staged changes, un-pushed changes, or diffs between the current and working branch, rendering findings as suggestions, comments, or LGTMs inside a read-only *magit-hutch: code review* buffer. Hutch builds on Magit and Transient, lets users queue suggestions with `m` and bulk-apply them with `A` scope-aware to the reviewed files, and defaults to reviewing staged changes only.

read9 min views4 publishedOct 6, 2026
hutch: local code reviews in emacs for the mildly disenfranchised
Image: source

I haven't had a real job in four years. I closed down a startup I'd been building, just last month. During all these years, I spent most of that time at the back end of the frontier of AI agents. But I've finally caught up. It's been some 500 days since coding agents have really picked up, and they're genuinely more productive than, previously, instructed.

Even though I still prefer the pedagogical aspect of AI over the task-completing automaton aspects, the latter has driven all sorts of tooling around reviewing code, and not just writing and deploying it. The typical review agent party-line is: agents jump in, before your colleagues do, spray logorrhea across twenty pull requests before you have had a chance to wake up and look at your phone. This works, sometimes, for some people. But if you're like me, you still have humans reviewing code before it ships to users, and it's better to respect those people and their time. This is the case, regardless of where you sit on the balance of game-changer to curmudgeon.

All that is to say, no matter which direction agents take to get better with time, I hope we still care about things. Not in the way of formalizing care, with high-fidelity agent instructions and prompts or some superior upholding of taste sort of thing, but something as simple as announcing: hey I'm still here, and I understand all this.

So as a long-time emacs user, I present yet another attempt at wedging LLMs, agents and coding harnesses, now inside your text buffers (!) with Hutch. It's a small, Magit-induced code-review interface that fits a standard Magit commit-push workflow locally and hopefully helps reclaim some load created upstream.

quick tour# #

Open up Magit, and hit the dispatcher binding (usually d) and you'll see a Hutch code review action put up next to the DWIM binding. Hutch operates on three different scopes: staged changes, un-pushed changes, and changes between current branch and working branch. By default, it's staged changes only, since that's most useful.

Once a review starts, you'll see a nice little progress bar in a new *magit-hutch: code review* buffer until the findings<sup>1</sup> are complete. This is a read-only buffer, but you can still perform the pre-bound actions.

Each finding has a type (suggestion, comment or LGTM), a file name, relevant line numbers, a title and a description. Suggestions additionally have a patch diff. Hutch piggybacks on Magit and Transient to render these menus so it behaves much like its own interface; keyboard-driven sub-menus, diff coloring, and highlighting. Suggestions are special since they can be applied. To mark a suggestion for application, you queue it with m.

Then bulk-apply all queued suggestions with A. The application is scope-aware, so if you queue a finding for staged changes, it will apply the fix directly to the staged files.

That's it! Getting started should hopefully be pretty simple and intuitive for existing emacs users. There are of course a few interesting things going on behind the scenes, some of which I'll cover in the next few sections.

patches over comments# #

A big UX handicap of showing review comments and suggested patches in-buffer is that there is no existing connective tissue of a commenting system. With GitHub, though, the review UI collapses outdated comments on new commits and most review bots sit over the suggestion mechanic if they have changes to suggest.

Hutch is made with a bias towards patches, rather than just prosaic comments. According to the Aider leaderboard (and through some of my own experiments), the SEARCH/REPLACE diffs are a lot more obedient across different models than just asking the model to author correct patches with precise line numbers.

For an Aider-style diff, you have to ensure there's enough surrounding context for the SEARCH to be unique, and ideally also preserve indentation. In Hutch's case, the tool's function schema naturally decomposes the file, search, and replace fields:

src/utils.clj
<<<<<<< SEARCH
(defn add [a b]
(defn add [a b c]
  (+ a b c))
>>>>>>> REPLACE

I've noticed that a lot of older (or cheaper) models tend to recall the SEARCH block from memory when asked for diffs, instead of copying it verbatim from read_file, read_diff or surrounding_context calls. This invariably botches them entirely. So we get them verified before submission. If SEARCH is missing or matches more than once, the finding is downgraded to a plain comment. On a unique hit, Hutch locally creates a unified diff:

Once a series of udiffs and comments are rendered, they can be marked and bulk applied. Hutch applies them per-file, lowest hunk first (bottom-up) so line positions are minimally disturbed. Each finding runs its own git apply and a bad application marks itself invalid so the rest can continue to land.

All this patching and commenting infrastructure pulls its weight, since with only a couple of keystrokes, you hopefully get less reading and parsing work and more actionable triaging. None of this guarantees patches-always of course, and it shouldn't.

With more powerful models, a simpler diffing method might generally work pretty well. But for a tool that's built to work across different and cheaper models, it's essential to be maximally supportive. In general, I feel like a key point of much of the agentic infrastructure we build is to have knobs for optimizing token:cost ratios. This could often mean thorny workarounds for good-enough models.

barely enough tooling# #

Hutch has a fairly minimal toolset for pulling context:

  1. read_diff
  2. read_file
  3. search_codebase
  4. surrounding_context

Out of these, surrounding_context is the more interesting one. It wraps over Tree-sitter and uses grammars that are installed. It works by letting the model widen out to the enclosing definition of a relevant line and further out, as needed. In my tests, the overall read token consumption compared to simply blasting read_file was anecdotally lower with comparable levels of review quality<sup>2</sup>.

All the findings from the model are submitted to the agent at once. On the write side of things, verify_block locally verifies diffs, and along with other comments and LGTM notices, submits them through a submit_review tool call. submit_review itself runs through some post-processing work, like gating hallucinations about files and line numbers, trimming the length of descriptions and downgrading patches to comments if they don't apply cleanly.

Once the submission lands, the output from all this work is persisted durably under refs/hutch/id and can be separately committed as a means of sharing (with magit-post-commit-hook) or for repainting later. If you squint hard enough, it might appear like a change identifier for a stacked-diff review tool, but its purpose is to keep reviews in the git tree, rather than identify changesets for human reviews. We don't really care about multi-party human reviews, it's all local.

evaluating# #

The one unfortunate part about benchmarking Hutch is how ungainly it is to pull comparison-ready output from text buffers. I initially ran the evals by invoking multiple headless emacsen and tee-ing the agent output before it was rendered, but eventually settled on emitting Perfetto traces and using them as the underlying medium for evals.

I haven't seen agents traced through Perfetto elsewhere. This is likely for good reason. They aren't meant for this kind of thing really. They don't have a first-class notion of what a "prompt" or a "tool call" is. It's designed for kernels and browsers and not an abstract system with tons of prose.

But for a single-player, emacs-local agent, it sort of works. You can answer all kinds of structural questions like why did this review take 40 rounds?, what tools were run in parallel?, or how much wall time was spent in reading diffs?, and so on. But more importantly, it's free and infra-free. If you set hutch-trace-dir, it will emit Perfetto traces and you can just load them up on ui.perfetto.dev. Easy.

Here's an example to fetch tool calls and their total times. This is the entire pipeline. No dashboards or SDKs required:

SELECT
  name                       AS tool,
  COUNT(*)                   AS calls,
  ROUND(SUM(dur) / 1e6, 1)   AS total_ms
FROM slice
WHERE category = 'tool'
GROUP BY name
ORDER BY total_ms DESC;

-- tool                 calls  total_ms
--------------------------------------
-- search_codebase      18     4210.3
-- read_diff            7      2103.1
-- surrounding_context  12     880.5
-- read_file            3      412.7
-- submit_review        1      42.9

With this set up, we take a mix of strategies from Martian’s code review benchmark and the CR-Bench preprint and compute Precision, Recall, and Fβ scores. The evals are described in more detail<sup>3</sup> in the eval/README.org section. But broadly, we run the bench against 40 PRs, 132 goldens, and use GPT 5.2 as a classifying judge. The eval pipeline goes off and runs queries directly on the traces. Looking at the numbers, I believe we land somewhere around the #16 mark on Martian’s Offline Benchmark leaderboard, which is pretty competitive for a no-memory, single-shot agent.

Outside of classified scoring, there are a few interesting things about the agent itself:

Different models tend to catch different bugs. Out of 132 goldens, each model hits 40-50 goldens, with an overlap of 18 hits across all three models. Which means hypothetically, if all three ran combined, it would catch ~55% more bugs than one model alone.

Pretty lousy agreement across the models on what a bug is, I'd say.

GPT 5.5 tends to hit my default round limit (80) a lot more than the other models for roughly the same hit rate. Opus 4.8 takes 3x fewer turns to complete.

On token efficiency, Opus is much cheaper on output tokens used per good finding by a respectable margin, but burns 3x more context on inputs, possibly due to the growing context Hutch resends each round.

dead on arrival# #

This is all probably too late, as I've been told. No one really writes or reviews code, uses editors or version control by hand any longer. I made this for myself and for workflows that I still practice. I don't want to purport any arguments about whether one should or shouldn't use LLMs with emacs. The tool has more to do with unlocking a certain kind of workflow than the overreach of agents in niche locations.

If this continues to be useful, I'd like to add a conversational mode for every finding (like CodeRabbit) and perhaps maintain a context tree learnt from and committable to the codebase to improve review quality and speed.

  1. In the example, I use GLM-5.2 as the underlying model, but this is configurable to whatever backend the excellent gptel project supports.↩
  2. The characterization tests and evals are covered under the evaluating section, but I haven't yet gotten a chance to verify this claim empirically.↩
  3. There are some biases and nuances to consider before treating the hard metrics as truly objective. But I've elided them from the post since they are described in more detail in the README .↩
── more in #ai-tools 4 stories · sorted by recency
── more on @hutch 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/hutch-local-code-rev…] indexed:0 read:9min 2026-10-06 · —