Your Vector Search Knows “Bank” Is Related to “Bank.” It Doesn’t Know Which Bank You Mean. A developer built ARBITER, a deterministic measurement engine that sits between vector retrieval and LLM generation to disambiguate word senses that similarity search conflates. In a benchmark compressing 768-dimensional representations to 72 dimensions, ARBITER retained 0.9653 similarity versus PCA's 0.8693 while separating ambiguous senses far more sharply (0.066 vs ~0.85 for "bank"), and it ranks candidate fields by coherence without generating answers. Vector search is very good at similarity. Similarity is not always the same thing as meaning. That distinction starts becoming expensive when retrieval feeds an LLM. Consider: financial bank river bank They share the same word. A similarity system has every reason to place them near each other. A useful reasoning system often needs to do the opposite. It needs to separate the senses. That problem is one of the reasons I built ARBITER . ARBITER is a deterministic measurement engine. You give it: context + a field of possibilities and it returns a coherence-ordered field. It does not generate an answer. It measures the possibilities you supplied. A common RAG pipeline looks roughly like: query ↓ vector retrieval ↓ top N chunks ↓ LLM The retrieval stage is intentionally broad. That is useful, but it also means bad context can survive long enough to reach generation. Once incorrect-but-related context enters the prompt, the generator has to reason around it. A different pipeline is: query ↓ vector retrieval ↓ candidate field ↓ ARBITER ↓ coherence-ordered field ↓ LLM ARBITER does not replace retrieval. It gives you a deterministic measurement step between retrieval and generation. Here is a live ARBITER call: curl -sS -X POST https://arbiter.grip.fyi/v1/compare \ -H 'content-type: application/json' \ --data '{ "query":"python memory", "candidates": "garbage collection", "malloc", "snake habitat" , "top k":3 }' The resulting ordering: 0.483573 garbage collection 0.324794 malloc 0.167992 snake habitat Same interface: state / intent / context + field of possibilities → ARBITER → ranked resonance field The field could contain: documents tools routes robot actions suppliers code paths hypotheses agents products The primitive does not change. :chatgpt-content-reference{index="0"} An earlier ARBITER compression/disambiguation benchmark compared a 768-dimensional source representation compressed to 72 dimensions using PCA versus ARBITER. The similarity-retention result was: PCA 0.8693 ARBITER 0.9653 But the more interesting result was sense separation. For ambiguous words: PCA ARBITER bank ~0.85 0.066 bat ~0.85 0.073 Lower here means better separation between competing senses. So river bank and financial bank remained strongly entangled after PCA compression, while ARBITER separated them much more sharply. The same benchmark reduced the representation from 768 dimensions to 72: a 10.7× dimensional reduction. :chatgpt-content-reference{index="1"} That is the part I care about. Not simply: Can I preserve similarity? But: Can the representation preserve enough structure to distinguish what something means in context? The same behavior shows up in ordinary ambiguous language. For: Best bass fishing spots in freshwater lakes ARBITER produced: 0.772 Largemouth bass in shallow weedy areas 0.542 Bass amplifiers and speaker impedance 0.293 Bass clef instruments in orchestra 0.272 Bass guitar string gauges Crane safety regulations on construction sites it produced: 0.828 Tower cranes require certified operators 0.325 Sandhill cranes migrate through Nebraska 0.274 Origami cranes symbolize peace in Japan 0.173 Crane flies are harmless insects And: Cell division rates in tumor growth analysis returned: 0.823 Mitotic cell division in tumor tissue 0.312 Prison cell division protocols 0.287 Cellular network division coverage These are not generated answers. They are measurements over an explicit candidate field. :chatgpt-content-reference{index="2"} This is where things get more interesting. Take Python . Without extra context: Programming 0.796 Snakes 0.284 Now change the supplied perspective: "As a herpetologist..." and the same meanings reorganize: Snakes 0.700 Programming 0.422 Apple behaves similarly: baseline: Tech company 0.861 Fruit 0.252 With: "As a chef..." the ordering flips: Fruit 0.668 Tech company 0.397 No retraining. The supplied context changed, so the field changed. :chatgpt-content-reference{index="3"} A lot of current AI infrastructure treats representation as a lookup problem: Which stored object is closest? But many useful machine decisions are closer to: Given this exact state, which of these possibilities fits best? Those are not identical questions. RAG is an obvious place to use that distinction because retrieval already gives you a bounded field. But the same operation applies to agent routing, tool selection, robotics, screening, planning, and other systems where the candidates already exist. The generator does not always need to make the decision. Sometimes the candidates are already there. What you need is a measurement. The live endpoint is: POST https://arbiter.grip.fyi/v1/compare Or install the lightweight CLI: curl -fsSL https://arbiter.grip.fyi/install | sh Then: arb "python memory" \ "garbage collection" \ "malloc" \ "snake habitat" The CLI is just the interface to the hosted ARBITER service. The current developer surface includes 10 successful calls per day free, after which the same endpoint moves to native x402 payment at $0.01 per call. :chatgpt-content-reference{index="4"} Try a field where you already know what the answer should be. Ambiguous words are a good place to start. ARBITER: https://arbiter.grip.fyi https://arbiter.grip.fyi Description: Similarity is not the same thing as meaning. A deterministic measurement step for RAG, reranking, and bounded decision fields. Tags: ai, rag, machinelearning, programming