cd /news/ai-agents/reconnected-claude-code-over-sse-and… · home › topics › ai-agents › article
[ARTICLE · art-144148] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Reconnected Claude Code Over SSE — and Found the Distance Scores Aren't as Stable as I Thought

A developer reconnected Claude Code to an MCP search server over SSE after switching it from stdio, then found that embedding distance scores swing by as much as 54 points for the same target document when a query is merely reworded, while identical query text returned the same score (437.7) across different clients and transports. The developer also noted that a seemingly accurate answer quoting a wrong command came from Claude Code's own filesystem grep rather than the RAG tool, whose snippets are truncated to about 200 characters, and that every query returned the same five results from a small index.

by read3 min views1 publishedOct 2, 2026

Loose end from Entry 14: switching mcp_search_server.py from stdio to SSE broke the Entry 08 Claude Code connection, which was registered expecting a spawned process, not a network server. Fixing it turned out to be the easy part:

claude mcp remove today-i-ran-notes
claude mcp add --transport sse today-i-ran-notes http://localhost:8090/sse
claude mcp list

✔ Connected, server still running from Entry 14. Clean.

Re-ran Entry 08's original test — ask it to use the tool, search for the oc pod-status question. It worked, and gave a solid narrative answer identifying Entry 02's wrong command. But I wanted the actual distance score this time, not just the summary, so I asked for the raw tool output directly. That's where this got more interesting than "confirmed the reconnect works."

Two calls came back, for two slightly different phrasings of basically the same question:

Query Distance to 02-oc-cli-mentor-system-prompt.md
"how to check pod status with oc" 430.1
"oc get pods command output troubleshooting" 375.8

Same target document, same intent, a 54-point swing just from rewording. For comparison, the gap between the best and worst result within a single query was 442.6 → 461.1 — about 18 points. The variance from paraphrasing the question was three times larger than the variance between a genuinely-relevant result and a marginal one in the same result set.

That's a sharper version of the open question from Entry 13, and it resolves cleanly once you see both numbers side by side. Entry 13 guessed the embedding model might be rewarding lexical/syntactic overlap over pure meaning — a broken, command-shaped query scoring tighter than a clean one. This data points at something more basic underneath that: distance isn't calibrated for "is this relevant" at all, full stop. It's only meaningful for ranking results within one specific query's wording. Compare distances across different phrasings of the same question, and the number tells you more about word choice than about relevance.

There's also a reproducibility result worth stating plainly, since it cuts the other way: Entry 14's n8n test and this one both queried the same exact string — "how do I check pod status with oc" — through different clients, and both landed on 437.7. That's not in tension with today's finding. It's the other half of it: identical query text reliably gives identical distance, regardless of client or transport. Reword the query, even slightly, and the number moves more than you'd expect from meaning alone.

One more thing worth being honest about, since it affects how much to trust the narrative answer from the first run: Claude Code's accurate quote of the wrong oc command — the one from Entry 02 — didn't actually come from the MCP tool. The tool's results are truncated to ~200 characters per snippet, and that quote wasn't in any of them. The session separately grepped and read the full file directly off disk. So the earlier clean-looking answer wasn't a pure test of the RAG pipeline; it was the RAG pipeline plus Claude Code's own filesystem access filling the gap. Worth knowing before holding that result up as proof the retrieval system alone is doing the work — some of it was something else entirely.

Collection size matters here too, and it's worth naming rather than letting the clean distance numbers imply otherwise: every query in this entry returned the exact same five results, just reordered. The index doesn't have much in it yet. These rankings are real, but they're ranking a very short list.

── more in #ai-agents 4 stories · sorted by recency
── more on @claude code 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/reconnected-claude-c…] indexed:0 read:3min 2026-10-02 · —