# IDE-Style Code Navigation and Refactoring Don’t Automatically Improve Your Coding Agent

> Source: <https://dev.to/b2a48b/ide-style-code-navigation-and-refactoring-dont-automatically-improve-your-coding-agent-28n>
> Published: 2026-10-06 14:36:13+00:00

The appeal of a drop-in upgrade for a coding agent is easy to understand: install a tool, give the agent IDE-style code navigation and refactoring, and expect better results. I wanted to know whether that actually helped the same agent finish tasks faster, with fewer mistakes or less token usage.

I tested one recommendation on ten refactorings in HTTPX. Each task had two attempts with the agent's ordinary tools and two with added navigation and refactoring tools: 40 completed runs.

Both workflows passed every acceptance check. With the added tools and explicit usage guidance, average agent time increased from **89.8 seconds to 111.4 seconds**—about **24% longer**. Average total reported tokens increased by about **40%**.

For this prescribed workflow in fresh sessions, the added tools increased time and token usage without improving acceptance. That is why I want reproducible evidence and tested best practices before treating an extra tool as a drop-in improvement.

The idea made sense: “go to definition” and “find references” help when working with unfamiliar code. Searching for a method name can return unrelated methods with the same name; finding references to a specific symbol can give a more precise answer.

The recommended tool was [Serena](https://github.com/oraios/serena), which exposes these capabilities to agents through MCP, a protocol for connecting tools to AI clients. An agent can retrieve a function, inspect its references, or edit a particular symbol.

One route is a language server, which supplies code analysis for editor features through the [Language Server Protocol (LSP)](https://microsoft.github.io/language-server-protocol/). A native IDE integration is another: the [JetBrains backend](https://github.com/oraios/serena#the-serena-jetbrains-plugin) uses an IDE's own analysis and refactoring capabilities.

This experiment used the language-server backend with Pyright, a Python analysis tool.

The practical question was whether adding those capabilities improved the same agent's finished work after accounting for the effort of using them.

Successful navigation and refactoring show capabilities. A productivity claim also needs evidence about completed work. For example, Serena's [agent-written evaluations](https://oraios.github.io/serena/04-evaluation/030_results/000_evaluation-results.html) compare operation chains, call counts, and payload sizes on its JetBrains backend. Those reports do not establish gains across repeated development tasks with independent acceptance checks.

Its [methodology](https://oraios.github.io/serena/04-evaluation/010_methodology.html#prompt-fairness) raises a confirmation-bias concern. The agent chooses examples and judges its own work. Negative findings are requested, yet the prompt used to review the methodology assumes capable agents use the tools correctly, excludes misuse and failure modes, and explicitly says: “Do not question the validity of this assumption.”

I would not treat reports built on that protected assumption as proof of a drop-in improvement. For any tool, I want repeated complete-task comparisons, independent checks, and tested guidance for real workflows.

I was looking for an existing, medium-sized Python project on GitHub: real code spread across files, enough variety for ten distinct refactorings, and a substantial test suite to check that behavior stayed intact. It also needed to be manageable for repeated runs.

[HTTPX](https://github.com/encode/httpx/tree/b5addb64f0161ff6bfe94c124ef76f6a1fba5254), a Python HTTP client, was an arbitrary choice that fit that brief. The selected snapshot had about 8,800 lines of production Python. Every attempt used the same fixed commit.

The ten tasks included renaming classes and methods, moving functions while preserving old imports, extracting shared implementations and modules, inlining a private helper, and simplifying control flow.

One task renamed `Headers.get_list` while leaving `QueryParams.get_list` unchanged. It also required preserving the old method as a compatibility alias. Another moved a file-length helper into a new module while preserving stream position and keeping the old import bound to the same function. These gave the checks something more specific to verify than whether the code still ran.

Every task had two attempts per condition, starting from fresh Docker copies of the same source. Runs were sequential, with the order reversed for the second repetition. The configuration was:

| Component | Tested configuration | 
|---|---|
| Agent | Codex CLI 0.160.0 | 
| Model | GPT-6.1 Sol, Medium reasoning | 
| Evaluation runner | Harbor 0.23.0 | 
| Serena | 1.7.0, LSP backend | 
| Python analysis | Pyright 1.1.403 | 
| Per-run resources | 2 CPUs, 4 GiB memory, Linux ARM64 Docker on a Mac | 

The ordinary workflow used shell commands, text search, file reads, and edits. The guided workflow kept those tools and added the tool server with explicit instructions: read its guidance, locate the target symbol, inspect references before editing, and use at least one applicable symbol-editing operation.

**This compares a prescribed workflow using the added tools with the agent's ordinary workflow.** The guidance is part of the intervention. It does not isolate the effect of LSP itself or measure how an agent chooses among optional tools.

Acceptance came from a separate offline verifier that checked change scope, requested structure, behavior, tests, type checking, and linting. Before the runs, it accepted ten known-good patches and rejected ten deliberately flawed ones. No human repaired the measured patches.

The full reproduction bundle is not public yet, so independent reproduction remains a limitation of this report.

| Measurement | Ordinary tools | Guided tools | Change | 
|---|---|---|---|
| Accepted attempts | 20/20 | 20/20 | Equal | 
| Mean agent time | 89.8 s | 111.4 s | +24% | 
| Mean total reported tokens | 147,247 | 205,444 | +40% | 
| Mean uncached input tokens | 20,508 | 30,047 | +47% | 
| Mean completed tool actions | 9.6 | 20.4 | +112% | 

For someone expecting a drop-in upgrade, this is a worse result on the measured efficiency criteria: the same accepted work took more time and tokens, despite explicit instructions to use the new tools.

All twenty attempts with the added tools followed the required workflow. Five of the 239 server calls failed.

Token figures are cumulative usage across model calls and include cached context. Billing was not measured.

Agent time includes MCP startup, tool use, and tests run by the agent. It excludes container and agent setup, and the independent verifier. Each attempt started a fresh tool server and language-server process, so these are results for fresh sessions rather than a server kept running throughout a working day.

When I compared matching attempts at the same task, the guided workflow was faster in **5 of 20 comparisons**. The table below averages each task's two attempts per workflow:

| Refactoring | Ordinary tools | Guided tools | 
|---|---|---|
| Class rename across files | 80.4 s | 92.3 s | 
| Rename one of two same-named methods | 76.9 s | 82.5 s | 
| Function move with compatibility | 104.4 s | 133.6 s | 
| Extract two classes into a module | 80.8 s | 122.1 s | 
| Extract related helpers into a module | 109.9 s | 173.0 s | 
| Extract a duplicate implementation | 73.2 s | 107.0 s | 
| Extract a shared encoder | 77.9 s | 88.2 s | 
| Extract a branch with edge cases | 104.8 s | 100.2 s | 
| Inline a private helper | 94.7 s | 121.7 s | 
| Simplify local control flow | 95.2 s | 93.7 s | 

The predeclared target was at least a 15% speed improvement without lower acceptance. It was not met.

Two repetitions per task give a limited view of run-to-run variation. These are descriptive results for one agent and model on ten curated tasks. Other coding agents, models, and Python projects may behave differently.

For any tool presented as a coding-agent upgrade, I want the authors to define "better," measure it, and turn the findings into tested guidance. People looking for a drop-in solution should not have to discover every trade-off themselves.

For a claim about speed, reliability, or token efficiency, I would look for:

A tool may help on one task and add overhead on another. Publish both outcomes and the conditions behind them: repository size, model, backend, server lifetime, and instructions. If selective use is recommended, test that policy too.

Until that evidence is inspectable and reproducible, I do not find broad productivity claims trustworthy enough to base my workflow on.

This experiment measured guided use on ten tasks with fresh tool servers. A larger repository, another language, a server that stays running, or a native IDE backend might produce a different result; those possibilities remain untested here. With both workflows passing every check, this comparison also cannot establish differences in maintainability or defects outside the tested cases.

For this workflow, the extra capabilities produced a slower result with more token usage. That is the risk of treating new tooling as an automatic upgrade: without reproducible evidence and tested best practices, you can make the same agent's workflow worse.

My question for the authors of coding-agent tools is straightforward: **can you demonstrate an improvement in complete tasks, publish enough of the experiment for someone else to verify it, and show users when and how to get that benefit?**
