How to use the AutoResearch plugin to let Claude Code test and revise its own coding instructions, with a full cost and setup breakdown.
What is AutoResearch for Claude Code? #
AutoResearch is a plugin that lets a coding agent run its own experiment loop: propose a change, test it against a fixed metric, keep the change if it scores better, discard it if it doesn’t, and repeat. Applied to Claude Code, this means you can hand Claude a skill file full of instructions (how to fix a bug, how to review a PR, how to write tests) and let it revise those instructions automatically based on measured results rather than guesswork. The underlying Claude model doesn’t change. The workflow wrapped around it does.
TL;DR #
- AutoResearch originated as a model-training loop built by Andrej Karpathy, where an agent edits training code, runs an experiment, checks a fixed metric, and keeps or discards the change based on the result.
- Udit Gulati’s community plugin ports that loop to Claude Code , letting you define a goal, an editable scope, a metric, and a verification command instead of running actual model training.
- The experiment works by optimizing a skill file , a markdown file holding the instructions Claude follows for a task like bug fixing, while keeping the skill’s name and grading rules fixed.
- A trustworthy result requires a frozen evaluator and held-out tasks that the optimizer never sees, otherwise the loop can learn to game its own benchmark instead of genuinely improving.
- Every trial needs to start from a clean, identical state using Claude’s print mode with a bare flag so you’re comparing instruction content, not inherited memory, plugins, or stale context.
- The cost adds up fast , with a small three-task, three-attempt setup producing dozens of separate Claude API calls before you even reach the held-out comparison.
- The payoff is narrow and specific , a verified improvement on your test tasks, not proof of a universally smarter coding agent.
How does the AutoResearch loop actually work? #
The core pattern has four pieces: a goal, an editable scope, a metric, and a verification command. You tell the plugin what you’re trying to improve (for example, “solve more of these bug fixing tasks under the same limits”), which files it’s allowed to touch (a single skill file, nothing else), how success is measured (percentage of attempts that pass fixed tests), and what command checks that score.
AutoResearch then proposes a candidate revision to the instructions, runs it through your verification command, records whether the score improved, and either keeps the new version or rolls back to the previous best. Each iteration produces one focused change plus a logged result, so you end up with a history showing which ideas helped and which ones didn’t.
This only works if “better” has a concrete definition. Asking an agent to “get better at coding” gives it nothing to check against. Asking it to “pass more of these three frozen bug-fix tests without touching forbidden files” gives it something measurable.
How do you set up a bug-fixing experiment in Claude Code? #
A basic setup needs Claude Code, git, Python (for a small evaluation helper), and Node.js (for the plugin’s hooks). The practical steps:
- Start in a clean, separate experiment folder with its own git repo, isolated from any real project work.
- Install the plugin inside Claude Code using
/plugin marketplace addwith the AutoResearch repository, then/plugin install autoresearch@autoresearch. Start a fresh session afterward so the plugin’s reference files load correctly. - Create a skill folder at
.claude/skills/bugfix/containing askill.mdfile with a name, description, and starting instructions. This is the file the optimizer will revise. Keep a separate saved copy of the original version and note its git commit before any changes, so you have something to compare against later. - Build an evaluation helper. AutoResearch doesn’t ship a coding benchmark, so you need your own: a broken project fixture, grading tests, and a results log. The worker (Claude attempting the fix) can only touch the disposable project. The optimizer can only touch the instructions. The grading tests stay fixed and untouchable by either.
- Create a development set of small, well-understood bugs. Three examples used in testing this approach: a duplicate-removal function that needs to preserve order, a function that incorrectly treats zero as a missing value, and a function that accidentally mutates the caller’s input list. Keep a couple of different tasks held out entirely from the optimizer’s feedback loop, reserved for final evaluation only.
Why does the evaluator need to be tested before you spend money? #
#
Plans first. Then code.
Remy writes the spec, manages the build, and ships the app.
Before running any real Claude trials, you need to confirm the measuring tool actually measures what you think it measures. That means feeding the evaluator a known broken implementation and confirming it fails, feeding it a correct patch and confirming it passes, and feeding it a patch that deletes a test or edits a forbidden file and confirming it gets rejected. You also want to check what happens on a timeout or authentication failure. An incomplete run should never quietly register as a pass.
Skipping this step risks the worst kind of failure: a loop that looks like it’s improving because the grading is broken, not because the instructions are better.
Why run trials with a bare flag and injected instructions? #
If every trial session inherits different memory, plugins, or leftover conversation context, you’re no longer comparing instruction files cleanly. The fix is running Claude in print mode with a bare flag, which skips automatic discovery of skills, plugins, memory, and project configuration files. The workflow instructions are then explicitly injected using an append-system-prompt-file flag, pulled straight from the skill file’s instruction body. This ensures both the baseline and the candidate versions get injected identically, so what’s being measured is the instruction content itself, not whether Claude happens to stumble on the right skill from a cluttered list.
One practical catch: this bare print mode doesn’t run through a normal Claude subscription login. It requires an API key or equivalent provider credentials, and those calls bill through that path. If you only have subscription access, you can approximate the comparison manually in fresh sessions, but it should be described as a less controlled version of the experiment, not treated as equivalent.
How do you know if the “improved” instructions actually improved anything? #
This is the part most likely to go wrong. A development score can rise because the revised instructions have become overly specific to the exact three example bugs, maybe referencing a particular function name or quietly encoding the expected fix. That produces an impressive-looking score with zero benefit to tomorrow’s actual bug.
The guard against this is a held-out test set the optimizer never sees during development. After freezing the best candidate from the iteration loop, run both the original and optimized instructions against the held-out tasks, using fresh project copies and fresh Claude sessions, with multiple attempts per task and the run order alternated so neither version always goes first. Once you look at a held-out failure and then revise instructions in response, that task has effectively become part of development and can’t be used to validate the result anymore.
Three outcomes are possible. The candidate improves on both development and held-out tasks, which is real evidence of a transferable gain on this task set (though not proof of a universal coding upgrade). Development improves but held-out performance stays flat or drops, meaning the loop overfit to its examples or the gain was noise. Or both versions solve everything, meaning the benchmark has hit a ceiling and you need a harder test to tell them apart.
Is running this loop worth the cost? #
Not automatically. A small setup with three development tasks and three attempts each produces nine trials per evaluated version. Running a baseline plus three candidate iterations adds up to 36 trial runs. Comparing the original and optimized versions on two held-out tasks, three attempts each, adds another 12. That’s 48 separate Claude coding trials before counting the optimizer’s own reasoning overhead or any manual rehearsal runs.
Remy is new. The platform isn't. #
Remy is the latest expression of years of platform work. Not a hastily wrapped LLM.
Claude’s print mode supports setting a maximum dollar budget and a maximum number of turns per run, and pairing those with an external timeout and a controller that halts further trials once a budget is hit is a sensible guardrail. The plugin’s scope and iteration settings are instructions given to the agent, not a hard filesystem sandbox or a guaranteed spending cap, so the actual cost controls need to live outside the loop itself.
Frequently Asked Questions #
What is the AutoResearch plugin for Claude Code?
It’s a community-built plugin (by Udit Gulati) that adapts Andrej Karpathy’s original experiment-loop concept, originally built for model training, into a general-purpose loop for Claude Code. It lets you define a goal, an editable file scope, a metric, and a verification command, then runs iterative candidate revisions against that metric.
Do you need a GPU to use AutoResearch with Claude Code?
No. The original AutoResearch project was built around training a model on an NVIDIA GPU with a short training budget per experiment. Using it to optimize Claude Code’s instructions borrows the experimental pattern (propose, test, keep or discard) without running any model training or needing GPU hardware.
Can this loop make Claude Code generally smarter?
No. The underlying Claude model never changes, only the instructions or workflow around it. Any measured improvement applies specifically to the tasks and metric you tested against. A handful of held-out bug-fix tasks passing more often is useful evidence, but it doesn’t establish a universal coding upgrade.
Why does the experiment need held-out tasks separate from development tasks?
Because the optimizer can inadvertently tune instructions to fit the specific examples it sees during development, sometimes by effectively encoding the answer rather than a general strategy. Held-out tasks that the optimizer never influences are the only way to check whether an improvement actually transfers to new problems.
How expensive is it to run a full AutoResearch experiment on Claude Code?
It scales quickly. A modest setup with three development tasks, three attempts per task, three candidate iterations, and a two-task held-out comparison run three times each can produce around 48 separate Claude trial runs, not counting the optimizer’s own overhead. Setting dollar budgets, turn limits, and external timeouts before running is strongly advisable.