Show HN: Find stale, orphaned, deleted-but-retrievable RAG vectors RAGproof released an open-source tool, rag-staleness-check, that performs read-only staleness, orphan, duplicate, and retrievability checks on a single pgvector, Qdrant, or Chroma index. The tool, installable via pip for Python 3.10+, reports findings to stdout and a JSON file, skipping checks with missing inputs rather than showing misleading zeros. It is the free half of RAGproof's paid audit, which cross-validates findings against a git-history-derived ledger across all three engines. Read-only staleness / orphan / duplicate / retrievability checks for a single pgvector, Qdrant or Chroma index - the open-source half of the RAGproof https://ragproof.io "decayed RAG index" teardown full writeup + multi-engine ledger-verified methodology https://ragproof.io/blog/rag-index-decay/ . This tool runs against your own already-indexed vector database and tells you: staleness - indexed chunks whose source document has since changed needs a manifest with last modified per document and the engine itself storing a per-row last-modified value orphans - indexed chunks whose source document no longer exists in your manifest at all duplicates - near-identical chunks cosine similarity ≥ threshold, default 0.98 , plus an exact-hash pass if you store a per-chunk content hash retrievable-after-delete - a basic probe: for ids you believe are deleted, is the vector still physically retrievable by id storage-layer persistence , and does it still surface in top-k search results functional leak ? It's read-only - it never calls a write/delete/upsert method on your engine. See Read-only guarantee read-only-guarantee below for how that's enforced, and its limits. This is the single-engine, ledger-free, self-serve slice. It doesn't include: - multi-engine orchestration run it once per engine yourself - the GDPR Article-17 reporting pack signed evidence, verifiable-erasure report - engine-specific deep checks pgvector dead-tuple/VACUUM detail, Qdrant optimizer-threshold detail, Chroma on-disk HNSW growth - a ledger-based precision/recall harness - there's no ground truth here, this tool reports what it finds, not how accurate the finding is Those live in the private, paid RAGproof https://ragproof.io audit, which cross-validates every finding against a git-history-derived ground-truth ledger across all three engines simultaneously. pip install rag-staleness-check pgvector only pip install "rag-staleness-check qdrant " + Qdrant pip install "rag-staleness-check chroma " + Chroma pip install "rag-staleness-check all " everything Or, without installing: pipx run rag-staleness-check --engine pgvector --dsn "$DSN" --source ./docs manifest.json Requires Python 3.10+. rag-staleness-check \ --engine pgvector \ --dsn "postgresql://readonly user:pw@localhost:5432/mydb" \ --pg-table chunks \ --pg-doc-id-column doc id \ --pg-last-modified-column last modified \ --pg-content-hash-column content sha256 \ --source ./docs manifest.json \ --deleted-ids ./deleted ids.json \ --out findings.json This prints a scorecard to stdout and writes the full result to --out default findings.json . If a check is missing a required input - no --source , no per-row metadata configured, no --deleted-ids - it reports "skipped": true with a reason instead of silently showing a misleadingly clean 0%. | Engine | --dsn format | Example | |---|---|---| pgvector | Postgres DSN | postgresql://user:pw@localhost:5432/db | qdrant | Base URL | http://localhost:6333 | chroma | host:port | localhost:8000 | Your table/column or payload/metadata field names are your own - this tool doesn't assume anything about your schema beyond what you tell it via flags: pgvector --pg- : --pg-table required , --pg-id-column default id , --pg-vector-column default embedding , --pg-doc-id-column , --pg-last-modified-column , --pg-content-hash-column . The last three are optional, and leaving them out is exactly what makes staleness / orphans / the duplicates check's exact-hash pass report skipped / zero-coverage. Qdrant / Chroma --collection required, plus payload/metadata field names : --doc-id-field default doc id , --last-modified-field default last modified , --content-hash-field default content sha256 . A JSON file describing what documents should exist - this is the join key for staleness and orphans. There's no ledger in this tool; the manifest is the only source of truth about "what's current." { "manifest version": 1, "documents": { "doc id": "handbook/engineering/on-call.md", "last modified": "2026-06-01T00:00:00Z", "content sha256": "optional, doc-level, informational only" } } doc id has to match whatever value your own ingestion pipeline wrote into each indexed chunk's doc-id column/field see schema mapping above . last modified is only needed if you want the staleness check to run. content sha256 here is doc-level and purely informational - it isn't the input to duplicate detection see below . Full example at examples/docs manifest.json https://github.com/rimironenko/rag-staleness-check/blob/v0.1.1/examples/docs manifest.json . A JSON array of ids you believe are deleted: "chunk-abc123", "chunk-def456" For each one, this checks whether the vector is still retrievable by id storage-layer persistence and, if so, whether it still shows up in a top-k self-query functional leak . The distinction matters: per EDPB Guidelines 05/2019 https://www.edpb.europa.eu/documents/guideline/guidelines-52019-on-the-criteria-of-the-right-to-be-forgotten-in-the-search en , erasure has to be verifiable and irreversible - suppressing a record from search results isn't enough on its own if the underlying vector is still there. findings.json 's retrievable.framing field spells this out every run. This tool never calls a write/delete/upsert method on your engine, and that's enforced two ways: Structurally, via a static test -walks the actual syntax tree of every file under tests/test no write methods.py src/rag staleness check/ and fails the build if any write/delete/ upsert-shaped method gets called, or any SQL write statement is passed to execute / executemany . Not a substring grep - that would false-positive on this very README and on docstrings - it parses real syntax. At runtime, for pgvector - every connection opens with SET default transaction read only = on , and rag staleness check.readonly.assert pg read only re-checks SHOW default transaction read only before any query runs, refusing to proceed if it isn't 'on' . That's defense-in-depth against a connection pooler e.g. PgBouncer silently swallowing the session-level SET . Qdrant and Chroma don't give a client-exposed way to check "is this session/API key read-only" - that's deployment-side RBAC, not something either client library reports. For those two engines, enforcement is the static test above plus your own deployment-side scoping: connect with a read-only-scoped API key/token if your engine supports one. There's no runtime assertion for Qdrant/Chroma equivalent to pgvector's - flagging that here rather than implying parity where there isn't any. For extra assurance on pgvector, connect through a locked-down role instead of relying on this tool alone: CREATE ROLE rag staleness check readonly NOSUPERUSER LOGIN PASSWORD '...'; GRANT CONNECT ON DATABASE mydb TO rag staleness check readonly; GRANT USAGE ON SCHEMA public TO rag staleness check readonly; GRANT SELECT ON chunks TO rag staleness check readonly; No default telemetry. Nothing is collected or sent automatically. --share-anonymous-scorecard is an explicit opt-in that prints the exact anonymized payload aggregate counts + engine type only - never your dsn, hostnames, doc ids, or chunk ids that would be shared. There's no backend configured yet, so nothing actually goes over the network - the flag is reserved for a future opt-in submission endpoint. DO NOT TRACK in the environment, if set, forces this off even if the flag is passed. Separately, this tool disables chromadb's own client-side telemetry anonymized telemetry=False when connecting to Chroma. The chromadb package ships its own default-on PostHog-based telemetry, independent of anything above - leaving it enabled would quietly break the "no default telemetry" promise for --engine chroma users even though this tool's own code never phones home. One cosmetic wrinkle: on chromadb==0.6.3 , the setting takes effect verified via client. system.settings.anonymized telemetry is False , but a startup event ClientStartEvent still tries to fire through a code path that doesn't consistently honor it, and throws capture takes 1 positional argument but 3 were given while constructing the call - before any request is made. You may see this line printed; it doesn't mean anything was sent. Duplicate detection's exact-hash pass needs a per-chunk content hash stored in your engine --pg-content-hash-column / --content-hash-field . Without one, it falls back to cosine-ANN only, and a warning field in the duplicates check says so rather than staying silent about it. Chroma's only reports a live count, intentionally snapshot stats - no on-disk HNSW-directory-size measurement. The private RAGproof audit's dev-environment build of this used a docker exec du trick against a known-local container; that has no equivalent against a real third-party deployment, so it isn't shipped here. - no on-disk HNSW-directory-size measurement. The private RAGproof audit's dev-environment build of this used a may warn if your installed client's minor version differs from your server's by more than 1 e.g. "Qdrant client version 1.18.0 is incompatible with server version 1.15.5" . Non-fatal - the check still runs - but if it's noisy, pin qdrant-client 's own compatibility check qdrant-client to match your server's minor version yourself. This tool reports counts and pairs, not precision/recall. There's no ground truth available client-side to score itself against - that's what the private, ledger-based audit is for. Duplicate detection needs a retrievable vector per candidate - a chunk with no vector returned by get vector is silently excluded from the cosine-ANN pass nothing to query with , not counted as "not a duplicate." qdrant-client and chromadb are optional extras pip install rag-staleness-check qdrant / chroma / all so a pgvector-only user doesn't have to pull in Chroma's heavier transitive dependencies.- Neither is pinned to an exact version. This tool connects to a third party's already-running deployment, whose server version is out of its control, so hard-pinning - unlike a project that ships and controls its own Docker images - would just break users on anything else. Apache-2.0. See LICENSE https://github.com/rimironenko/rag-staleness-check/blob/v0.1.1/LICENSE . Independent of any other RAGproof project's license. Issues and PRs welcome. Not yet implemented: a --method minhash surface-text-based duplicate detection mode via datasketch as an alternative/complement to the cosine-embedding approach above - PRs welcome.