# Show HN: CiteGuard – local CLI for checking citations in Markdown reports

> Source: <https://github.com/larrylot/citeguard>
> Published: 2026-09-07 05:59:14+00:00

**Local CLI that verifies citations in agent / deep-research Markdown reports.**

`citeguard check report.md` → per-citation verdicts: URL resolve, title/host soft-match, optional claim–source overlap.

- **No API keys** for core checks
- **No telemetry** , no SaaS, no account
- **Offline `--fixtures` mode** for CI
- MIT licensed

Exploration bet for The Lord (0 SEK). Not a merchant product.

|  | CiteGuard | Ask ChatGPT / another LLM | 
|---|---|---|
| Reproducible | Deterministic resolve + scores | Non-deterministic prose | 
| Offline CI | `--fixtures` planted cases | Needs network + model | 
| Cost / keys | Free local HTTP | API key / subscription | 
| Audit trail | JSON verdicts per URL | Chat transcript | 
| Hallucination check on the checker | No LLM in the loop | Can invent “looks fine” | 

CiteGuard does **not** claim to prove a source supports a legal/scientific conclusion. It flags **dead links, title bait, weak overlap, and shady redirects** — the failure modes that show up in agent research dumps.

Portable `SKILL.md` (Cursor / Claude / Codex / agentskills.io format):

- Canonical: [`skill/SKILL.md`](/larrylot/citeguard/blob/main/skill/SKILL.md)
- Cursor mirror: [`.cursor/skills/citeguard/SKILL.md`](/larrylot/citeguard/blob/main/.cursor/skills/citeguard/SKILL.md)
- skills CLI layout: [`skills/citeguard/SKILL.md`](/larrylot/citeguard/blob/main/skills/citeguard/SKILL.md)

```
npx skills add larrylot/citeguard -s citeguard -y
cd citeguard
python3 -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
# Offline CI-style check on a clean planted report
citeguard check fixtures/clean.md --fixtures

# A bad report (dead links) — exits 1
citeguard check fixtures/dead_link.md --fixtures

# Title-bait SEO farm planted as "asyncio docs"
citeguard check fixtures/title_mismatch.md --fixtures

# Machine-readable
citeguard check fixtures/mixed.md --fixtures --json
```

Example human output:

```
CiteGuard — fixtures/mixed.md
Summary: {'clean': 3, 'dead': 1, 'title_mismatch': 1, 'total': 5}
------------------------------------------------------------
[OK] L3 https://docs.github.com/en/actions
...
[DEAD] L7 https://example.invalid/dead-citation-404
...
[TITLE] L9 https://example.invalid/python-asyncio-guide
```

Live network check (no fixtures):

```
citeguard check path/to/agent-report.md
citeguard check path/to/agent-report.md --json --no-overlap
```

| Verdict | Meaning | 
|---|---|
| `clean` | Resolved; title/claim checks passed | 
| `dead` | 404 or network failure | 
| `http_error` | Non-success HTTP (e.g. 500) | 
| `title_mismatch` | Link text soft-match vs `<title>` too low | 
| `claim_weak` | Claim sentence tokens barely appear in page text | 
| `redirect_suspect` | Cross-host redirect without strong title match | 
| `unresolved` | URL missing from fixtures catalog (fixtures mode only) | 

Synthetic agent-style Markdown under [`examples/realworld/`](/larrylot/citeguard/blob/main/examples/realworld) (dead links, DNS failures, title bait on `example.com`, plus working RFC / Example Domain controls).

- Findings table (FACT counts): [`examples/realworld/REPORT.md`](/larrylot/citeguard/blob/main/examples/realworld/REPORT.md)
- Per-dump JSON + text: [`examples/realworld/out/`](/larrylot/citeguard/blob/main/examples/realworld/out)
- Live demo page: [https://larrylot.github.io/citeguard/](https://larrylot.github.io/citeguard/)

[`corpus/`](/larrylot/citeguard/blob/main/corpus) — 20 short public-domain-style fake agent-research Markdown snippets with planted failures (`example.invalid`, title bait via `example.com` / `httpbin.org`). Documented for benchmarks. See [` corpus/README.md`](/larrylot/citeguard/blob/main/corpus/README.md).

One-page JSON output demo: [`docs/index.html`](/larrylot/citeguard/blob/main/docs/index.html) (GitHub Pages: [https://larrylot.github.io/citeguard/](https://larrylot.github.io/citeguard/)).

- [CiteGuard vs alternatives](/larrylot/citeguard/blob/main/docs/vs-alternatives.md) — honest comparison vs “ask ChatGPT”, LinkChecker, html-proofer, ReportBench (FACT / ASSUMPTION labeled)

```
docker build -t citeguard .
docker run --rm -v "$PWD":/data -w /data citeguard check fixtures/mixed.md --fixtures
```

`fixtures/` ships **9** Markdown reports + `*.expected.json` + HTML pages + `catalog.json` for offline resolve:

- clean cites, dead links, title mismatch, claim–source mismatch, redirect suspect, HTTP error, footnotes, mixed, bare URLs

```
pytest -q
# accuracy gate: ≥80% resolve+classify on fixtures
```

Free form — **HUMAN SETUP** ≤10 min: see [`WAITLIST.md`](/larrylot/citeguard/blob/main/WAITLIST.md). Paste the public URL below when ready:

Waitlist: `<!-- HUMAN: paste Tally or Google Form URL -->`

- [`SHOW_HN.md`](/larrylot/citeguard/blob/main/SHOW_HN.md) — Show HN draft
- [`POSTS.md`](/larrylot/citeguard/blob/main/POSTS.md) — Reddit / LinkedIn templates

Agent does **not** post or outreach.

See [`STATUS.md`](/larrylot/citeguard/blob/main/STATUS.md). Kill if <10 stars **and** <10 waitlist after 7 days with ≥1 human public post, or fixture accuracy <80%, or HN reads it as a chatbot wrapper.

MIT — see [`LICENSE`](/larrylot/citeguard/blob/main/LICENSE).
