cd /news/artificial-intelligence/don-t-use-cosine-similarity-careless… Β· home β€Ί topics β€Ί artificial-intelligence β€Ί article
[ARTICLE Β· art-100237] src=github.com β†— pub= topic=artificial-intelligence verified=true sentiment=Β· neutral

Don't use cosine similarity carelessly We fixed it this way

VectorPrism, an open-source retrieval engine from insightitsGit, claims to cut multi-vector storage costs to a single 1024-dimensional vector per chunk by multiplexing six specialized relevance subspaces plus a 16-float control header, using HNSW for stage-1 and intent-gated rescoring for stage-2. In a calibrated hard_adversarial finance pack, dense retrieval achieved only 7.1% Recall@10 and missed 13 of 14 queries, while the multi-channel approach recovered all 13 missed queries at 100% (13/13) with ~1.8 ms latency at 1000-doc scale, though the authors caution the results are not generalizable to all corpora.

read11 min views2 publishedAug 17, 2026
Don't use cosine similarity carelessly We fixed it this way
Image: Michielbdejong (auto-discovered)

Positional Subspace Multiplexing (PSM) & Intent-Gated 2-Stage Retrieval Engine for High-Scale RAG.

One contiguous

1024dtensor. Six independently trained relevance subspaces. Stage-1 HNSW + Stage-2 intent-gated rescoring. Baseline vector-DB storage cost β€” not 6Γ— multi-vector inflation.

** Interactive demo** Β·

Β·

BenchmarksΒ·

Pilot guide

Technical reportKeywords: pgvector multi-vector cost reduction, Intent-gated RAG retrieval engine, Causal retrieval for enterprise RAG, Positional subspace multiplexing vector search, Reduce hallucinations in root-cause RAG, VectorPrism, HNSW

Scope honesty: the table below is from our calibrated hard_adversarial

finance pack (dense is designed to miss). It is not a claim that every public corpus shows the same Miss@10. For partner corpora use scripts/corpus_recovery_audit.py β€” see

.

demos/external_audit/

Dense fails on purpose in that pack. Multi-channel recovers those misses.

Metric Result
Dense R@10 7.1%
Dense Miss@10 13/14 (93%)
Multi z-score recovered@10 13/13 (100%)
RRF recovered@10 (conservative) 10–11/13 (77–85%)
Auto-graph recovered@10 11/13 (85%)
1000-doc scale recovered@10 13/13 (~1.8 ms)

Full tables, caveats, and reproduce commands β†’ BENCHMARKS.md

Interactive query comparison (dense vs multi) β†’ demo site

Raw JSON/MD artifacts β†’ demos/finance_demo/results/

Enterprise RAG is stuck between two bad defaults:

Flat cosine over a single embeddingβ€” semantically β€œclose” neighbors that are causally wrong, taxonomically wrong, or temporally expired. Teams call themfunny neighbors; production calls themhallucination fuel.** Multi-vector indexing**(one ANN index per representation) β€” better signal, but** 500%–1,000%**storage and query fan-out on pgvector / Qdrant bills.

VectorPrism multiplexes six specialized representation subspaces plus a 16-float Control Header into a single 1024-dimensional contiguous buffer per chunk:

Constraint VectorPrism answer
Storage 1Γ— vector footprint (one vector(1024) / named full tensor)
Stage 1 HNSW only on the 368d dense core slice
Stage 2 In-RAM zero-copy slice scoring with intent weights
Early exit Header filters (epistemic_truth , anchor_dist , model_version ) before heavy math
Latency target < 15ms end-to-end search SLA (see benchmarks)

Philosophically grounded channel design. Engineering-grounded memory contract. Production path for pgvector and Qdrant.

Keywords: causal retrieval, incident log RAG, DevOps root-cause analysis, β€œwhy did the service fail”

When on-call asks β€œWhy did Server X crash at 3 AM?”, cosine-only RAG returns symptom-adjacent text. VectorPrism’s Time ODE & Directional Causality slice ([896:1024)

) is trained with an asymmetric bilinear score (q^{\top} M c) (PSMRetrievalEngine.causal_score

). Intent routing up-weights the causal channel on β€œwhy / cause / reason” queries so Stage-2 rescoring prefers causeβ†’effect order, not merely lexical neighbors.

Keywords: hyperbolic embeddings RAG, taxonomy search, medical ontology retrieval, legal hierarchy search

Parent–child trees distort badly in Euclidean space. The Hyperbolic Taxonomy (Porphyry) slice ([640:768)

) lives in a PoincarΓ© ball (norm < 1) and is scored with PoincarΓ© distance in Stage 2. Hierarchy intents (β€œcategory”, β€œparent”, β€œtype of”, β€œtree”) shift IntentClassifier

weights toward hyperbolic structure for medical, legal, and product taxonomies.

Keywords: bitemporal retrieval, compliance RAG, healthcare audit trail, finance document expiry filter

Before Stage-2 matrix math, the 16d Control Header Manifest ([0:16)

) exposes O(1) metadata:

  • Epistemic truth score (soft by default; hard filter opt-in after ECE calibration)
  • Identity anchor distance (OOD / injection-risk gate in Stage 1)
  • Exact int64 timestamp (packed in the header for audit / future filters β€” not applied as a Stage-1 SQL/Qdrant predicate today) - Model version for safe re-ingest after retrains ( applied as a Stage-1 filter; search defaults to the checkpoint’smodel_version

)

Stage 1 rejects low-truth, high-anchor-distance, or wrong-model_version

chunks before rescoring.

Keywords: multi-vector RAG cost reduction, pgvector HNSW, Qdrant named vectors, high-scale vector search

Instead of six ANN indexes, VectorPrism stores one 1024d tensor. Stage 1 indexes only the generated 368d dense_core_slice. Stage 2 pulls the full tensor for the top-~100 candidates and rescored slices in RAM. AI SaaS platforms keep multi-signal retrieval without multi-vector sticker shock.

Ground truth: PSMTensorContract

/ VectorPrismTensorContract

in tensor_contract.py.

1024-d VectorPrism Tensor (float32)
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ [  0 ..  15]  16d   Control Header Manifest                              β”‚
β”‚ [ 16 .. 383] 368d   Dense Semantic Core          (Hume / Wittgenstein)   β”‚
β”‚ [384 .. 511] 128d   Relational Group Algebra     (Aristotle / Al-Khwarizmi)β”‚
β”‚ [512 .. 639] 128d   Disentangled Latent Space    (Jabir)                 β”‚
β”‚ [640 .. 767] 128d   Hyperbolic Taxonomy          (Porphyry)              β”‚
β”‚ [768 .. 895] 128d   Identity Consistency         (Ibn Sina)              β”‚
β”‚ [896 ..1023] 128d   Time ODE & Causality         (Mulla Sadra / Spinoza) β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
         β–² Stage-1 HNSW indexes ONLY dense_core [16:384) β†’ 368 dims

| Inclusive range | Code slice (start:end ) | Dims | Channel | Role | |---|---|---|---|---| [0000..0015] | HEADER [0:16) | 16 | Control Header Manifest | Bitmask, truth, anchor dist, timestamp, model version | [0016..0383] | DENSE_CORE [16:384) | 368 | Dense Semantic Core | L2-normalized cosine space; Stage-1 ANN | [0384..0511] | RELATIONAL [384:512) | 128 | Relational Group Algebra | Train: TransE (S+R\approx O); serve today: L2 proximity (-|q_{\mathrm{rel}}-c_{\mathrm{rel}}|) (no query-time relation id yet) | [0512..0639] | DISENTANGLED [512:640) | 128 | Disentangled Latent (Jabir) | VIB latent (z) | [0640..0767] | HYPERBOLIC [640:768) | 128 | Hyperbolic Taxonomy (Porphyry) | PoincarΓ© ball | [0768..0895] | IDENTITY [768:896) | 128 | Identity Consistency (Ibn Sina) | Distance-to-frozen (v_0); Stage-1 gate only | [0896..1023] | CAUSAL_TIME [896:1024) | 128 | Time ODE & Causality (Spinoza) | Scored as (q^{\top} M c) |

Header sub-layout (exact packing via PSMTensorContract.pack_header

/ unpack_header

):

Slot Field Encoding
[0]
Channel bitmask uint32 ↔ float32 bit reinterpret
[1]
Epistemic truth float32 in [0, 1]
[2]
Identity anchor distance float32
[3:5]
Bitemporal timestamp int64 ↔ 2Γ—float32 bit reinterpret
[5]
Model version uint32 ↔ float32 bit reinterpret
[6:16]
Reserved zero-filled
pip install "vectorprism[all]"

git clone https://github.com/insightitsGit/VectorPrism.git
cd VectorPrism
python -m venv .venv && source .venv/bin/activate   # Windows: .venv\Scripts\activate
pip install -U pip
pip install -e ".[all]"

vectorprism version
vectorprism pilot-check
pytest test_psm.py test_phases.py -q

The PyPI wheel ships schema.sql

and data/*.example.jsonl

(enough for pilot-check

/ run-all-smoke

). Full adversarial finance packs stay in git, not on PyPI.

Publish / release: PUBLISH.md Β· External pilot:

Β· Production:

PILOT.md

PRODUCTION.md

Core deps: torch

, numpy

, scipy

, scikit-learn

. Optional extras: encoder

, postgres

, qdrant

, dev

, all

.

docker compose up -d db
docker compose run --rm test
docker compose run --rm finance-pg
docker compose run --rm production-smoke
  • DB: localhost:5433

Β· DSNpostgresql://vectorprism:vectorprism@localhost:5433/vectorprism

  • Results: demos/finance_demo/results/

(PRODUCTION_RESULTS.md

, eval, live search JSON) - Full checklist: Β· Docker notes:PRODUCTION.md

DOCKER.md

Encode raw text with a frozen 768d encoder β†’ MultiTaskProjectionAdapter

β†’ contiguous 1024d tensor (matches ingestion_adapter.py

  • ingest_pipeline.py

).

import time
import torch
import numpy as np

from base_encoder import SentenceTransformerEncoder
from ingestion_adapter import MultiTaskProjectionAdapter, VectorPrismProjectionAdapter
from tensor_contract import PSMTensorContract as C, VectorPrismTensorContract
from losses import anchor_distance_score

encoder = SentenceTransformerEncoder("sentence-transformers/all-mpnet-base-v2")
adapter = MultiTaskProjectionAdapter(base_dim=768)  # alias: VectorPrismProjectionAdapter
adapter.eval()

texts = ["Cache eviction storm preceded the 3 AM outage on Server X."]
base = encoder.encode(texts)  # (1, 768)

header = C.pack_header(
    bitmask=C.default_channel_bitmask({"dense": True, "identity": True, "causal": True}),
    epistemic_truth=1.0,
    anchor_distance=0.0,
    timestamp=int(time.time()),
    model_version=1,
)
header_t = torch.from_numpy(header).unsqueeze(0)  # (1, 16)

with torch.no_grad():
    tensor_1024d, raw = adapter(base, header_t)
    dist = anchor_distance_score(raw["identity"], adapter.identity_anchor_v0)
    out = tensor_1024d.cpu().numpy().astype(np.float32)
    out[0, C.HDR_ANCHOR.start] = float(dist[0].item())

assert out.shape == (1, 1024)
assert out[0, C.DENSE_CORE.start:C.DENSE_CORE.end].shape == (368,)
meta = C.unpack_header(out[0])
print(meta)  # bitmask, epistemic_truth, anchor_distance, timestamp, model_version

Production shorthand (upsert path):

from checkpointing import load_checkpoint
from db_client import PgVectorClient  # or QdrantVectorClient
from ingest_pipeline import VectorPrismIngestPipeline, IngestDocument
from base_encoder import SentenceTransformerEncoder

ckpt = load_checkpoint("checkpoints/vectorprism.pt")
encoder = SentenceTransformerEncoder("sentence-transformers/all-mpnet-base-v2")
db = PgVectorClient("postgresql://user:pass@localhost:5432/vectorprism")

pipe = VectorPrismIngestPipeline(
    encoder=encoder,
    adapter=ckpt["adapter"],
    db=db,
    model_version=ckpt["model_version"],
    enabled_channels=ckpt.get("enabled_channels"),
)
pipe.upsert_documents([
    IngestDocument(document_id="inc-42", chunk_text="Cache eviction preceded the outage."),
])

Intent classification β†’ HNSW on dense core β†’ zero-copy slice rescoring (PSMRetrievalEngine.search

).

from checkpointing import load_checkpoint
from base_encoder import SentenceTransformerEncoder
from db_client import PgVectorClient
from ingest_pipeline import VectorPrismIngestPipeline
from retrieval_engine import PSMRetrievalEngine, IntentClassifier, VectorPrismRetrievalEngine
from tensor_contract import PSMTensorContract as C

ckpt = load_checkpoint("checkpoints/vectorprism.pt")
encoder = SentenceTransformerEncoder("sentence-transformers/all-mpnet-base-v2")
db = PgVectorClient("postgresql://user:pass@localhost:5432/vectorprism")

pipe = VectorPrismIngestPipeline(encoder, ckpt["adapter"], db, model_version=ckpt["model_version"])
engine = PSMRetrievalEngine(  # alias: VectorPrismRetrievalEngine
    db_client=db,
    causal_matrix=ckpt["causal_matrix"],  # learned M for qα΅€ M c
    hard_truth_filter=False,              # keep soft until ECE-calibrated
)

query_text = "Why did Server X crash at 3 AM?"
query_1024d = pipe.encode_query(query_text)

w_intent, filters = engine.classifier.classify(query_text)
print("w_intent=", w_intent, "filters=", filters)

hits = engine.search(query_1024d, query_text, top_k=5)
for h in hits:
    print(h["document_id"], h["final_score"], h.get("chunk_text", "")[:120])

CLI equivalents:

python train.py --channel dense --data data/dense_pairs.example.jsonl \
  --encoder sentence-transformers/all-mpnet-base-v2 --out checkpoints/vectorprism.pt

python vectorprism.py ingest --checkpoint checkpoints/vectorprism.pt \
  --documents data/documents.example.jsonl --backend pgvector --dsn "$VECTORPRISM_PG_DSN"

python vectorprism.py search --checkpoint checkpoints/vectorprism.pt \
  --query "Why did Server X crash at 3 AM?" --backend pgvector --dsn "$VECTORPRISM_PG_DSN"

Exact DDL from schema.sql:

CREATE EXTENSION IF NOT EXISTS vector;

CREATE TABLE IF NOT EXISTS psm_document_embeddings (
    id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
    document_id VARCHAR(255) NOT NULL UNIQUE,
    chunk_text TEXT NOT NULL,

    -- Full 1024-Dimensional Composite Tensor Payload
    tensor_1024d vector(1024) NOT NULL,

    -- Generated Column for Stage 1 Dense Core Slice [16..383] (368d)
    -- pgvector subvector() is 1-indexed: (17, 368) == zero-indexed [16:384)
    dense_core_slice vector(368) GENERATED ALWAYS AS (
        subvector(tensor_1024d, 17, 368)
    ) STORED,

    epistemic_truth FLOAT NOT NULL DEFAULT 1.0,
    anchor_dist FLOAT NOT NULL DEFAULT 0.0,
    valid_timestamp BIGINT NOT NULL,
    model_version INTEGER NOT NULL DEFAULT 0,

    created_at TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP
);

CREATE INDEX IF NOT EXISTS idx_psm_dense_core_hnsw
ON psm_document_embeddings
USING hnsw (dense_core_slice vector_cosine_ops)
WITH (m = 16, ef_construction = 64);

CREATE INDEX IF NOT EXISTS idx_psm_epistemic_truth ON psm_document_embeddings (epistemic_truth);
CREATE INDEX IF NOT EXISTS idx_psm_anchor_dist ON psm_document_embeddings (anchor_dist);
CREATE INDEX IF NOT EXISTS idx_psm_model_version ON psm_document_embeddings (model_version);
CREATE INDEX IF NOT EXISTS idx_psm_document_id ON psm_document_embeddings (document_id);

Apply:

psql "$VECTORPRISM_PG_DSN" -f schema.sql

Matches QdrantVectorClient

in db_client.py:

from qdrant_client import QdrantClient
from qdrant_client.http import models as qmodels

client = QdrantClient(url="http://localhost:6333")
collection = "psm_document_embeddings"

if not client.collection_exists(collection):
    client.create_collection(
        collection_name=collection,
        vectors_config={
            "dense_core_slice": qmodels.VectorParams(
                size=368,
                distance=qmodels.Distance.COSINE,
                hnsw_config=qmodels.HnswConfigDiff(m=16, ef_construct=128),
            ),
            "full_tensor": qmodels.VectorParams(
                size=1024,
                distance=qmodels.Distance.COSINE,
                hnsw_config=qmodels.HnswConfigDiff(m=0),
            ),
        },
    )

Payload fields used for Stage-1 filters: epistemic_truth

, anchor_dist

, model_version

. Payload also stores valid_timestamp

, chunk_text

, and document_id

(timestamp is header metadata today; not a Stage-1 predicate).

Targets enforced by the design and benchmark_harness.py

/ live_benchmark.py

budgets:

Stage Operation Budget
Header Bitmask / header unpack [0:16) via PSMTensorContract.unpack_header
< 0.05 ms
Stage 1 HNSW coarse search on dense_core_slice (368d) + header filters β†’ top 100
< 10 ms
Stage 2 RAM zero-copy slice rescoring (z-score fuse Γ— w_intent )
< 2 ms
E2E
Encode path excluded in pure Stage-2 harness; search SLA
< 15 ms
python benchmark_harness.py

python vectorprism.py live-benchmark \
  --checkpoint checkpoints/vectorprism.pt \
  --documents data/documents.example.jsonl \
  --backend memory --n-trials 20 --p95-budget-ms 15

Architecture path (unchanged pillars):

Text ─► Frozen 768d Encoder ─► MultiTaskProjectionAdapter ─► 1024d tensor
                              β”‚
                              β–Ό
              pgvector / Qdrant (dense HNSW + full tensor)
                              β”‚
         IntentClassifier ─► w_intent + filters
                              β”‚
         Stage 1: HNSW(dense_core_slice) + truth/anchor filters
                              β”‚
         Stage 2: dense / rel / dis / hyp / causal scores β†’ top-k
                  (Identity is Stage-1 gate only β€” not double-counted)

Channels are earned, not assumed. One channel at a time:

python train.py --channel dense --data your_pairs.jsonl \
  --encoder sentence-transformers/all-mpnet-base-v2 --out checkpoints/vectorprism.pt

python vectorprism.py eval --checkpoint checkpoints/vectorprism.pt \
  --documents your_docs.jsonl --eval your_eval.jsonl \
  --encoder sentence-transformers/all-mpnet-base-v2

python train.py --channel causal --data your_causal.jsonl \
  --init checkpoints/vectorprism.pt --out checkpoints/vectorprism.pt

See IMPLEMENTATION_SPEC.md for phased Definitions of Done and the gap matrix.

Building a regulated RAG stack, a multi-tenant AI SaaS retrieval plane, or a private compliance-aware knowledge system?

Insight ITS works with enterprise architects on:

  • Custom multi-task adapter fine-tuning for your ontology / incident / audit corpora
  • Private compliance connectors (bitemporal filters, calibrated epistemic truth, HITL review)
  • Managed control planes for versioned re-ingest across pgvector & Qdrant fleets

Talk to us

Apache License 2.0 β€” see LICENSE (or repository license metadata).

If VectorPrism informs your research or production retrieval stack:

@software{vectorprism2026,
  title  = {VectorPrism: Positional Subspace Multiplexing for Intent-Gated Retrieval},
  author = {Amin Parva},
  year   = {2026},
  url    = {https://github.com/insightitsGit/VectorPrism}
}

VectorPrism β€” six signals, one tensor, baseline storage cost, intent-gated speed.

── more in #artificial-intelligence 4 stories Β· sorted by recency
── more on @vectorprism 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/don-t-use-cosine-sim…] indexed:0 read:11min 2026-08-17 Β· β€”