cd /news/artificial-intelligence/designing-mcp-tools-for-a-7b-model-n… · home topics artificial-intelligence article
[ARTICLE · art-86155] src=dev.to ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Designing MCP Tools for a 7B Model, Not a 70B One

An engineer built an agentic assistant for battery engineering that pairs a physics-based digital twin with a local 7B language model, but found the small model struggled with multi-step tool orchestration, dropping tool calls in comparisons and producing incomplete answers. The project uses PyBaMM to simulate an LG M50 cell and wraps it as six MCP tools, with a LangGraph ReAct agent running on Ollama's qwen2.5:7b. The failure prompted a redesign of tools to reduce the planning burden on smaller models.

read6 min views2 publishedAug 4, 2026

I built an agentic assistant for battery engineering. A physics-based digital twin runs the actual electrochemical simulation, an LLM orchestrates and explains. Partway through, I ran into a failure that had nothing to do with the physics or the prompt wording. It was about how many tool calls I was asking a 7B model to plan in one go, and it changed how I designed every tool after that.

The principle behind the whole project: physics does the maths, the model does the language.

A PyBaMM digital twin simulates an LG M50 21700 cell using the open Chen2020 parameter set — a real electrochemical model, not a language model's best guess at a discharge curve. That twin is wrapped as six tools and exposed through FastMCP over stdio, so it's not hard-wired into one agent; any MCP client can drive it. On the other side, a LangGraph ReAct agent, backed by a local Ollama model, reads a plain-English question, decides which tool to call, and turns the physics result into an explanation.

The six tools:

Tool What it simulates
cell_info\
Static description of the modelled cell
simulate_discharge\
Constant-current discharge (runtime, delivered capacity/energy)
simulate_cccv_charge\
CC-CV charge (charge time, energy in)
compare_charging_strategies\
Runs and ranks multiple charge rates in one call
compare_discharge_rates\
Runs and compares multiple discharge rates in one call
simulate_degradation\
Capacity fade over N cycles via SEI growth

The model never touches a number. It picks a tool, reads the structured result, and narrates it. That split only holds up, though, if the agent reliably calls the right tool — and that's where a smaller local model started to show its limits.

Single-tool questions worked from early on. Ask "how long does it run at 1C?" and the agent calls simulate_discharge(c_rate=1.0)\

, reads the result, answers. Reliable every time.

The break came with comparisons. Ask "2C versus 0.5C" and the natural ReAct pattern is to call simulate_discharge\

twice, once per rate, then reason over both results. That's standard multi-step tool use, and it's the kind of thing larger models handle without issue. Running on qwen2.5:7b\

, though, multi-rate discharge comparisons dropped a rate: the model would not reliably chain two simulate_discharge\

calls, so the answer came back based on only one of the two rates, stated as if it covered the comparison the user actually asked for.

That's a quiet failure. It doesn't throw an error. It produces a confident, well-formatted answer that's simply incomplete, the exact failure mode the physics-grounding was supposed to prevent, just relocated one layer up, into tool orchestration instead of number generation.

It wasn't the only rough edge. Two others, for context: llama3.1:8b\

at one point printed tool calls as literal JSON text instead of actually invoking them, so no simulation ran at all, which is what pushed the model choice to qwen2.5:7b\

. And separately, the agent would sometimes speculate about why a number looked a certain way, labelling a delivered capacity above the 5.0 Ah nominal rating as "inefficiency" or "over-discharge," when a capacity above nominal at gentle rates is just normal cell behaviour, not something the simulation had reported as a problem.

It's tempting to read "the model dropped a tool call" as "get a bigger model" and stop there. That's true as far as it goes, but it skips the more useful question: what am I actually asking a 7B model to plan?

A two-call comparison isn't one decision, it's several in sequence — call tool A, hold its result in context, decide to call tool B with different arguments, hold that result too, then reconcile both before answering. Every one of those is a place a small model can drop a step, and on a 7B model run locally, without the depth of training data that gives larger models a stronger prior on "plan, then execute, then synthesize," that margin is thin.

Instead of asking the agent to orchestrate two simulate_discharge\

calls and reconcile them itself, I added compare_discharge_rates\

, mirroring the compare_charging_strategies\

tool that already existed for the charging side. It takes a list of rates, runs every simulation internally in Python, and returns runtime and capacity for all of them in one structured result.

The agent's job shrinks to one decision: recognize this is a comparison question, call the one tool built for it, narrate the result. No multi-step plan to hold in working memory, no partial result to silently drop. The system prompt makes the preference explicit too — it tells the model to prefer compare_discharge_rates\

and compare_charging_strategies\

over calling the single-simulation tools repeatedly, so the model isn't left to rediscover the right pattern on its own each time.

This is a small code change with a bigger implication for how I think about tool design for agents. The instinct, especially coming from examples built around large frontier models, is to expose small, composable primitives and let the agent chain them however a question requires. That's the right call when the model doing the chaining is strong. When it isn't, composability becomes a liability, because every extra hop is another place a plan can quietly fall apart. The tool boundary isn't just an API decision — it's a decision about how much sequential reasoning you're willing to hand to the model versus how much you do deterministically in code before the model ever sees it.

This didn't make the problem disappear everywhere, and I'd rather say that directly than imply otherwise. Chaining different tools in sequence — a charge comparison followed by a degradation run, with the agent reasoning across both — can still drop a step on a 7B model. The system prompt and the single-call comparison tools solve the specific case that came up most (comparing the same kind of simulation across several rates), not general multi-tool planning. For that, the honest options are a bigger model like qwen2.5:14b\

, or an explicit planning step ahead of execution, and I've left that as a known limitation rather than pretending the fix generalizes further than it does.

None of this is specific to batteries or to qwen2.5:7b\

. Any project running a smaller model locally — for cost, latency, privacy, or air-gap reasons — is going to hit some version of this. The generalizable move is: test the multi-step cases you expect people to actually ask for, watch for the quiet failures rather than just the crashes, and design tool boundaries around what the model actually does reliably, not around what looks cleanest in an architecture diagram. It's a less glamorous kind of engineering than swapping in a bigger model and moving on. It's also the kind that matters when the actual constraint is a system that runs entirely on hardware someone already owns, with no API bill and no dependency on a provider staying online.

If you're running a small model locally, where's the line you've found between letting the agent chain tool calls itself versus just collapsing the sequence into one call in code? Curious whether others have hit this with 7B-class models specifically, or whether it's more about the stdio/MCP setup than the model size. The full project — the PyBaMM twin, the MCP server, the LangGraph agent, and the rest of the issues hit along the way — is open source: github.com/BinuShefieldShifani/Batterytwin-mcp. Portfolio and other projects at binushefieldshifani.github.io.``

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @pybamm 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/designing-mcp-tools-…] indexed:0 read:6min 2026-08-04 ·