{"slug": "hutch-local-code-reviews-in-emacs-for-the-mildly-disenfranchised", "title": "hutch: local code reviews in emacs for the mildly disenfranchised", "summary": "Developer adjaecent released Hutch (magit-hutch), a local code-review interface for Emacs that runs LLM-powered reviews on staged changes, un-pushed changes, or diffs between the current and working branch, rendering findings as suggestions, comments, or LGTMs inside a read-only *magit-hutch: code review* buffer. Hutch builds on Magit and Transient, lets users queue suggestions with `m` and bulk-apply them with `A` scope-aware to the reviewed files, and defaults to reviewing staged changes only.", "body_md": "# hutch: local code reviews in emacs for the mildly disenfranchised\n\nI haven't had a real job in four years. I closed down a [startup](https://tramline.app) I'd been building, just last month. During all these years, I spent most of that time at the back end of the frontier of AI agents. But I've finally caught up. It's been some [500 days](https://en.wikipedia.org/wiki/List_of_large_language_models#2025) since coding agents have really picked up, and they're genuinely more productive than, previously, [instructed](https://www.youtube.com/watch?v=U_cSLPv34xk).\n\nEven though I still prefer the pedagogical aspect of AI over the task-completing automaton aspects, the latter has driven all sorts of tooling around reviewing code, and not just writing and deploying it. The typical review agent party-line is: agents jump in, before your colleagues do, spray logorrhea across twenty pull requests before you have had a chance to wake up and look at your phone. This works, sometimes, for some people. But if you're like me, you still have humans reviewing code before it ships to users, and it's better to respect those people and their time. This is the case, regardless of where you sit on the balance of game-changer to curmudgeon.\n\nAll that is to say, no matter which direction agents take to get better with time, I hope we still *care* about things. Not in the way of formalizing care, with high-fidelity agent instructions and prompts or some superior upholding of taste sort of thing, but something as simple as announcing: *hey I'm still here, and I understand all this*.\n\nSo as a long-time emacs user, I present yet another attempt at wedging LLMs, agents and coding harnesses, now inside your text buffers (!) with [Hutch](https://github.com/adjaecent/magit-hutch). It's a small, Magit-induced code-review interface that fits a standard Magit commit-push workflow *locally* and hopefully helps reclaim some load created upstream.\n\n## quick tour[#](#quick_tour)\n\nOpen up Magit, and hit the dispatcher binding (usually `d`) and you'll see a `Hutch code review` action put up next to the DWIM binding. Hutch operates on three different scopes: staged changes, un-pushed changes, and changes between current branch and working branch. By default, it's staged changes only, since that's most useful.\n\nOnce a review starts, you'll see a nice little progress bar in a new `*magit-hutch: code review*` buffer until the findings[<sup>1</sup>](#fn-1) are complete. This is a read-only buffer, but you can still perform the [pre-bound actions](https://github.com/adjaecent/magit-hutch#usage).\n\nEach finding has a type (suggestion, comment or LGTM), a file name, relevant line numbers, a title and a description. Suggestions additionally have a patch diff. Hutch piggybacks on [Magit](https://magit.vc) and [Transient](https://github.com/magit/transient) to render these menus so it behaves much like its own interface; keyboard-driven sub-menus, diff coloring, and highlighting. Suggestions are special since they can be applied. To mark a suggestion for application, you queue it with `m`.\n\nThen bulk-apply all queued suggestions with `A`. The application is scope-aware, so if you queue a finding for staged changes, it will apply the fix directly to the staged files.\n\nThat's it! Getting started should hopefully be pretty simple and intuitive for existing emacs users. There are of course a few interesting things going on behind the scenes, some of which I'll cover in the next few sections.\n\n## patches over comments[#](#patches_over_comments)\n\nA big UX handicap of showing review comments and suggested patches in-buffer is that there is no existing connective tissue of a commenting system. With GitHub, though, the review UI collapses outdated comments on new commits and most review bots sit over the [suggestion](https://docs.github.com/en/pull-requests/how-tos/review-pull-requests/incorporating-feedback-in-your-pull-request#applying-suggested-changes) mechanic if they have changes to suggest.\n\nHutch is made with a bias towards patches, rather than just prosaic comments. According to the [Aider leaderboard](https://aider.chat/docs/leaderboards) (and through some of my own experiments), the `SEARCH/REPLACE` diffs are a lot more obedient across different models than just asking the model to author correct patches with precise line numbers.\n\nFor an [Aider-style diff](https://aider.chat/docs/more/edit-formats.html), you have to ensure there's enough surrounding context for the `SEARCH` to be unique, and ideally also preserve indentation. In Hutch's case, the tool's function schema naturally decomposes the *file, search, and replace* fields:\n\n```\nsrc/utils.clj\n<<<<<<< SEARCH\n(defn add [a b]\n  (+ a b))\n=======\n(defn add [a b c]\n  (+ a b c))\n>>>>>>> REPLACE\n```\n\nI've noticed that a lot of older (or cheaper) models tend to recall the `SEARCH` block from memory when asked for diffs, instead of copying it verbatim from `read_file`, `read_diff` or `surrounding_context` calls. This invariably botches them entirely. So we get them verified before submission. If `SEARCH` is missing or matches more than once, the finding is downgraded to a plain comment. On a unique hit, Hutch locally creates a unified diff:\n\nOnce a series of udiffs and comments are rendered, they can be marked and bulk applied. Hutch applies them per-file, lowest hunk first (bottom-up) so line positions are minimally disturbed. Each finding runs its own `git apply` and a bad application marks itself `invalid` so the rest can continue to land.\n\nAll this patching and commenting infrastructure pulls its weight, since with only a couple of keystrokes, you hopefully get less reading and parsing work and more actionable triaging. None of this guarantees patches-always of course, and it shouldn't.\n\nWith more powerful models, a simpler diffing method might generally work pretty well. But for a tool that's built to work across different and cheaper models, it's essential to be maximally supportive. In general, I feel like a key point of much of the agentic infrastructure we build is to have knobs for optimizing token:cost ratios. This could often mean thorny workarounds for good-enough models.\n\n## barely enough tooling[#](#barely_enough_tooling)\n\nHutch has a fairly minimal toolset for pulling context:\n\n1. `read_diff`\n2. `read_file`\n3. `search_codebase`\n4. `surrounding_context`\n\nOut of these, `surrounding_context` is the more interesting one. It wraps over [Tree-sitter](https://batsov.com/articles/2026/02/27/building-emacs-major-modes-with-treesitter-lessons-learned/#why-tree-sitter) and uses grammars that are installed. It works by letting the model widen out to the enclosing definition of a relevant line and further out, as needed. In my tests, the overall read token consumption compared to simply blasting `read_file` was anecdotally lower with comparable levels of review quality[<sup>2</sup>](#fn-2).\n\nAll the findings from the model are submitted to the agent at once. On the write side of things, `verify_block` locally verifies diffs, and along with other comments and LGTM notices, submits them through a `submit_review` tool call. `submit_review` itself runs through some post-processing work, like gating hallucinations about files and line numbers, trimming the length of descriptions and downgrading patches to comments if they don't apply cleanly.\n\nOnce the submission lands, the output from all this work is persisted durably under `refs/hutch/id` and can be separately committed as a means of sharing (with `magit-post-commit-hook`) or for repainting later. If you squint hard enough, it might appear like a change identifier for a stacked-diff [review tool](https://blog.tangled.org/stacking), but its purpose is to keep reviews in the git tree, rather than identify changesets for human reviews. We don't really care about multi-party human reviews, it's all local.\n\n## evaluating[#](#evaluating)\n\nThe one unfortunate part about benchmarking Hutch is how ungainly it is to pull comparison-ready output from text buffers. I initially ran the evals by invoking multiple headless emacsen and tee-ing the agent output before it was rendered, but eventually settled on emitting [Perfetto](https://perfetto.dev) traces and using them as the underlying medium for evals.\n\nI haven't seen agents traced through Perfetto elsewhere. This is likely for good reason. They aren't meant for this kind of thing really. They don't have a first-class notion of what a \"prompt\" or a \"tool call\" is. It's designed for kernels and browsers and not an abstract system with tons of prose.\n\nBut for a single-player, emacs-local agent, it sort of works. You can answer all kinds of structural questions like *why did this review take 40 rounds?*, *what tools were run in parallel?*, or *how much wall time was spent in reading diffs?*, and so on. But more importantly, it's free and infra-free. If you set `hutch-trace-dir`, it will emit Perfetto traces and you can just load them up on [ui.perfetto.dev](https://ui.perfetto.dev). Easy.\n\nHere's an example to fetch tool calls and their total times. This is the entire pipeline. No dashboards or SDKs required:\n\n```\nSELECT\n  name                       AS tool,\n  COUNT(*)                   AS calls,\n  ROUND(SUM(dur) / 1e6, 1)   AS total_ms\nFROM slice\nWHERE category = 'tool'\nGROUP BY name\nORDER BY total_ms DESC;\n\n-- tool                 calls  total_ms\n--------------------------------------\n-- search_codebase      18     4210.3\n-- read_diff            7      2103.1\n-- surrounding_context  12     880.5\n-- read_file            3      412.7\n-- submit_review        1      42.9\n```\n\nWith this set up, we take a mix of strategies from [Martian’s code review](https://github.com/withmartian/code-review-benchmark) benchmark and the [CR-Bench preprint](https://arxiv.org/html/2603.11078v1) and compute Precision, Recall, and Fβ scores. The evals are described in more detail[<sup>3</sup>](#fn-3) in the [eval/README.org](https://github.com/adjaecent/magit-hutch/blob/main/eval/README.org) section. But broadly, we run the bench against 40 PRs, 132 goldens, and use GPT 5.2 as a classifying judge. The eval pipeline goes off and runs queries directly on the traces. Looking at the numbers, I believe we land somewhere around the #16 mark on Martian’s Offline Benchmark [leaderboard](https://codereview.withmartian.com/?mode=offline), which is pretty competitive for a no-memory, single-shot agent.\n\nOutside of classified scoring, there are a few interesting things about the agent itself:\n\nDifferent models tend to catch different bugs. Out of 132 goldens, each model hits 40-50 goldens, with an overlap of 18 hits across all three models. Which means hypothetically, if all three ran combined, it would catch ~55% more bugs than one model alone.\n\nPretty lousy agreement across the models on what a bug is, I'd say.\n\nGPT 5.5 tends to hit my default round limit (80) a lot more than the other models for roughly the same hit rate. Opus 4.8 takes 3x fewer turns to complete.\n\nOn token efficiency, Opus is much cheaper on output tokens used per good finding by a respectable margin, but burns 3x more context on inputs, possibly due to the growing context Hutch resends each round.\n\n## dead on arrival[#](#dead_on_arrival)\n\nThis is all probably too late, as I've been told. No one really writes or reviews code, uses editors or version control by hand any longer. I made this for myself and for workflows that I still practice. I don't want to purport any arguments about whether one should or shouldn't use LLMs with emacs. The tool has more to do with unlocking a certain kind of workflow than the overreach of agents in niche locations.\n\nIf this continues to be useful, I'd like to add a conversational mode for every finding (like CodeRabbit) and perhaps maintain a context tree learnt from and committable to the codebase to improve review quality and speed.\n\n1. In the example, I use GLM-5.2 as the underlying model, but this is configurable to whatever backend the excellent [gptel](https://github.com/karthink/gptel) project supports.[↩](#fnref1)\n2. The characterization tests and evals are covered under the [evaluating](#evaluating) section, but I haven't yet gotten a chance to verify this claim empirically.[↩](#fnref2)\n3. There are some biases and nuances to consider before treating the hard metrics as truly objective. But I've elided them from the post since they are described in more detail in the [README](https://github.com/adjaecent/magit-hutch/blob/main/eval/README.org) .[↩](#fnref3)", "url": "https://wpnews.pro/news/hutch-local-code-reviews-in-emacs-for-the-mildly-disenfranchised", "canonical_source": "https://kitallis.in/p/hutch-a-local-code-review-interface-for-magit/", "published_at": "2026-10-06 05:39:57+00:00", "updated_at": "2026-10-06 06:16:05.354057+00:00", "lang": "en", "topics": ["ai-tools", "developer-tools", "ai-agents", "large-language-models"], "entities": ["Hutch", "magit-hutch", "Emacs", "Magit", "Transient", "adjaecent", "GitHub", "Tramline"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/hutch-local-code-reviews-in-emacs-for-the-mildly-disenfranchised", "markdown": "https://wpnews.pro/news/hutch-local-code-reviews-in-emacs-for-the-mildly-disenfranchised.md", "text": "https://wpnews.pro/news/hutch-local-code-reviews-in-emacs-for-the-mildly-disenfranchised.txt", "jsonld": "https://wpnews.pro/news/hutch-local-code-reviews-in-emacs-for-the-mildly-disenfranchised.jsonld"}}