# Show HN: Find stale, orphaned, deleted-but-retrievable RAG vectors

> Source: <https://github.com/rimironenko/rag-staleness-check>
> Published: 2026-08-06 17:13:54+00:00

Read-only staleness / orphan / duplicate / retrievability checks for a
single pgvector, Qdrant or Chroma index - the open-source half of the
[RAGproof](https://ragproof.io) "decayed RAG index" teardown
([full writeup + multi-engine ledger-verified methodology](https://ragproof.io/blog/rag-index-decay/)).

This tool runs against **your own** already-indexed vector database and
tells you:

**staleness**- indexed chunks whose source document has since changed (needs a manifest with`last_modified`

per document and the engine itself storing a per-row last-modified value)**orphans**- indexed chunks whose source document no longer exists in your manifest at all** duplicates**- near-identical chunks (cosine similarity ≥ threshold, default 0.98), plus an exact-hash pass if you store a per-chunk content hash**retrievable-after-delete**- a basic probe: for ids you believe are deleted, is the vector still physically retrievable by id (storage-layer persistence), and does it still surface in top-k search results (functional leak)?

It's **read-only** - it never calls a write/delete/upsert method on your
engine. See [Read-only guarantee](#read-only-guarantee) below for how
that's enforced, and its limits.

This is the single-engine, ledger-free, self-serve slice. It doesn't include:

- multi-engine orchestration (run it once per engine yourself)
- the GDPR Article-17 reporting pack (signed evidence, verifiable-erasure report)
- engine-specific deep checks (pgvector dead-tuple/VACUUM detail, Qdrant optimizer-threshold detail, Chroma on-disk HNSW growth)
- a ledger-based precision/recall harness - there's no ground truth here, this tool reports what it finds, not how accurate the finding is

Those live in the private, paid [RAGproof](https://ragproof.io) audit,
which cross-validates every finding against a git-history-derived
ground-truth ledger across all three engines simultaneously.

```
pip install rag-staleness-check              # pgvector only
pip install "rag-staleness-check[qdrant]"     # + Qdrant
pip install "rag-staleness-check[chroma]"     # + Chroma
pip install "rag-staleness-check[all]"        # everything
```

Or, without installing:

```
pipx run rag-staleness-check --engine pgvector --dsn "$DSN" --source ./docs_manifest.json
```

Requires Python 3.10+.

```
rag-staleness-check \
  --engine pgvector \
  --dsn "postgresql://readonly_user:pw@localhost:5432/mydb" \
  --pg-table chunks \
  --pg-doc-id-column doc_id \
  --pg-last-modified-column last_modified \
  --pg-content-hash-column content_sha256 \
  --source ./docs_manifest.json \
  --deleted-ids ./deleted_ids.json \
  --out findings.json
```

This prints a scorecard to stdout and writes the full result to `--out`

(default `findings.json`

). If a check is missing a required input - no
`--source`

, no per-row metadata configured, no `--deleted-ids`

- it
reports `"skipped": true`

with a reason instead of silently showing a
misleadingly clean 0%.

| Engine | `--dsn` format |
Example |
|---|---|---|
`pgvector` |
Postgres DSN | `postgresql://user:pw@localhost:5432/db` |
`qdrant` |
Base URL | `http://localhost:6333` |
`chroma` |
`host:port` |
`localhost:8000` |

Your table/column or payload/metadata field names are your own - this tool doesn't assume anything about your schema beyond what you tell it via flags:

**pgvector** (`--pg-*`

): `--pg-table`

(required), `--pg-id-column`

(default `id`

), `--pg-vector-column`

(default `embedding`

),
`--pg-doc-id-column`

, `--pg-last-modified-column`

,
`--pg-content-hash-column`

. The last three are optional, and leaving them
out is exactly what makes staleness / orphans / the duplicates check's
exact-hash pass report `skipped`

/ zero-coverage.

**Qdrant / Chroma** (`--collection`

required, plus payload/metadata field
names): `--doc-id-field`

(default `doc_id`

), `--last-modified-field`

(default `last_modified`

), `--content-hash-field`

(default
`content_sha256`

).

A JSON file describing what documents *should* exist - this is the join
key for staleness and orphans. There's no ledger in this tool; the
manifest is the only source of truth about "what's current."

```
{
  "manifest_version": 1,
  "documents": [
    {
      "doc_id": "handbook/engineering/on-call.md",
      "last_modified": "2026-06-01T00:00:00Z",
      "content_sha256": "optional, doc-level, informational only"
    }
  ]
}
```

`doc_id`

has to match whatever value your own ingestion pipeline wrote
into each indexed chunk's doc-id column/field (see schema mapping above).
`last_modified`

is only needed if you want the staleness check to run.
`content_sha256`

here is doc-level and purely informational - it isn't
the input to duplicate detection (see below). Full example at
[ examples/docs_manifest.json](https://github.com/rimironenko/rag-staleness-check/blob/v0.1.1/examples/docs_manifest.json).

A JSON array of ids you believe are deleted:

```
["chunk-abc123", "chunk-def456"]
```

For each one, this checks whether the vector is still retrievable by id
(storage-layer persistence) and, if so, whether it still shows up in a
top-k self-query (functional leak). The distinction matters: per
[EDPB Guidelines 05/2019](https://www.edpb.europa.eu/documents/guideline/guidelines-52019-on-the-criteria-of-the-right-to-be-forgotten-in-the-search_en),
erasure has to be *verifiable and irreversible* - suppressing a record
from search results isn't enough on its own if the underlying vector is
still there. `findings.json`

's `retrievable.framing`

field spells this
out every run.

This tool never calls a write/delete/upsert method on your engine, and that's enforced two ways:

**Structurally, via a static test**-walks the actual syntax tree of every file under`tests/test_no_write_methods.py`

`src/rag_staleness_check/`

and fails the build if any write/delete/ upsert-shaped method gets called, or any SQL write statement is passed to`execute()`

/`executemany()`

. Not a substring grep - that would false-positive on this very README and on docstrings - it parses real syntax.**At runtime, for pgvector**- every connection opens with`SET default_transaction_read_only = on`

, and`rag_staleness_check.readonly.assert_pg_read_only()`

re-checks`SHOW default_transaction_read_only`

before any query runs, refusing to proceed if it isn't`'on'`

. That's defense-in-depth against a connection pooler (e.g. PgBouncer) silently swallowing the session-level`SET`

.

Qdrant and Chroma don't give a client-exposed way to check "is this session/API key read-only" - that's deployment-side RBAC, not something either client library reports. For those two engines, enforcement is the static test above plus your own deployment-side scoping: connect with a read-only-scoped API key/token if your engine supports one. There's no runtime assertion for Qdrant/Chroma equivalent to pgvector's - flagging that here rather than implying parity where there isn't any.

For extra assurance on pgvector, connect through a locked-down role instead of relying on this tool alone:

```
CREATE ROLE rag_staleness_check_readonly NOSUPERUSER LOGIN PASSWORD '...';
GRANT CONNECT ON DATABASE mydb TO rag_staleness_check_readonly;
GRANT USAGE ON SCHEMA public TO rag_staleness_check_readonly;
GRANT SELECT ON chunks TO rag_staleness_check_readonly;
```

**No default telemetry.** Nothing is collected or sent automatically.
`--share-anonymous-scorecard`

is an explicit opt-in that prints the exact
anonymized payload (aggregate counts + engine type only - never your dsn,
hostnames, doc_ids, or chunk_ids) that *would* be shared. There's no
backend configured yet, so nothing actually goes over the network - the
flag is reserved for a future opt-in submission endpoint.
`DO_NOT_TRACK`

in the environment, if set, forces this off even if the
flag is passed.

Separately, this tool disables **chromadb's own client-side telemetry**
(`anonymized_telemetry=False`

) when connecting to Chroma. The `chromadb`

package ships its own default-on PostHog-based telemetry, independent of
anything above - leaving it enabled would quietly break the "no default
telemetry" promise for `--engine chroma`

users even though this tool's
own code never phones home.

One cosmetic wrinkle: on `chromadb==0.6.3`

, the setting takes effect
(verified via `client._system.settings.anonymized_telemetry is False`

),
but a startup event (`ClientStartEvent`

) still tries to fire through a
code path that doesn't consistently honor it, and throws
`capture() takes 1 positional argument but 3 were given`

while
*constructing* the call - before any request is made. You may see this
line printed; it doesn't mean anything was sent.

**Duplicate detection's exact-hash pass** needs a per-chunk content hash stored in your engine (`--pg-content-hash-column`

/`--content-hash-field`

). Without one, it falls back to cosine-ANN only, and a`warning`

field in the`duplicates`

check says so rather than staying silent about it.**Chroma's** only reports a live count, intentionally`snapshot_stats()`

- no on-disk HNSW-directory-size measurement. (The private RAGproof
audit's dev-environment build of this used a
`docker exec du`

trick against a known-local container; that has no equivalent against a real third-party deployment, so it isn't shipped here.)

- no on-disk HNSW-directory-size measurement. (The private RAGproof
audit's dev-environment build of this used a
may warn if your installed client's minor version differs from your server's by more than 1 (e.g. "Qdrant client version 1.18.0 is incompatible with server version 1.15.5"). Non-fatal - the check still runs - but if it's noisy, pin`qdrant-client`

's own compatibility check`qdrant-client`

to match your server's minor version yourself.**This tool reports counts and pairs, not precision/recall.** There's no ground truth available client-side to score itself against - that's what the private, ledger-based audit is for.**Duplicate detection needs a retrievable vector per candidate**- a chunk with no vector returned by`get_vector`

is silently excluded from the cosine-ANN pass (nothing to query with), not counted as "not a duplicate."

`qdrant-client`

and`chromadb`

are**optional extras**(`pip install rag-staleness-check[qdrant]`

/`[chroma]`

/`[all]`

) so a pgvector-only user doesn't have to pull in Chroma's heavier transitive dependencies.- Neither is pinned to an exact version. This tool connects to a third party's already-running deployment, whose server version is out of its control, so hard-pinning - unlike a project that ships and controls its own Docker images - would just break users on anything else.

Apache-2.0. See
[LICENSE](https://github.com/rimironenko/rag-staleness-check/blob/v0.1.1/LICENSE).
Independent of any other RAGproof project's license.

Issues and PRs welcome. Not yet implemented: a `--method minhash`

surface-text-based duplicate detection mode (via `datasketch`

) as an
alternative/complement to the cosine-embedding approach above - PRs
welcome.
