My friend was working on project documentation where the README, configuration, results, and datasets could easily drift apart.
Instead of manually checking every number and configuration value, Receipts provides a quick, auditable report showing which claims are supported and which need attention.
Project folders can contain coursework, datasets, configuration files, and other information that shouldn't automatically be uploaded to a cloud AI provider.
Receipts uses Gemma locally through Ollama, so:
Most importantly:
The AI is not the final authority.
| Verdict | Meaning |
|---|---|
| VERIFIED | Evidence supports the claim. |
| CONFLICT | Evidence contradicts the claim. |
| AMBIGUOUS | Multiple plausible pieces of evidence exist. |
| UNVERIFIABLE | Suitable evidence could not be found. |
Each result keeps provenance so the user can understand where the claim came from and what evidence was used.
Receipts does not execute project code, notebooks, or scripts while inspecting a project.
Receipts generates a self-contained HTML verification report containing:
The controlled demo contains examples of all four verdict types:
VERIFIED · CONFLICT · AMBIGUOUS · UNVERIFIABLE
GitHub: https://github.com/sm4006/Receipts
The CLI generates a self-contained report.html file containing the verification results and provenance.
I used GitHub Copilot as a coding agent throughout development.
I worked from a written specification, implemented the system incrementally, reviewed each stage, added regression tests, and performed adversarial testing before moving to the next stage.
The main lesson from building this was that adding an LLM isn't enough.
You have to define exactly what the model is allowed to do.
For Receipts, that boundary is simple: That makes the result easier to test, audit, and trust.
Final test suite covered: Result: 85 tests → OK
The boundary between probabilistic AI and deterministic Python.
So:
AI → understand
Python → verify
Because when a project says:
"Our model achieved 94.2% accuracy."
I want the project to have the receipts — the evidence behind it.
Hardest part: deciding what the LLM was allowed to do.
Receipts was built for students/researchers who want a second check before submission.
Point it at the folder.
It checks.
If it can’t prove something, it says so. Receipts directly qualifies for the Hacktoberfest $100 partner categories:
Both categories emphasize open-source AI and Copilot-assisted development, which were central to how Receipts was built.
Extensions:
Principle stays the same:
AI extracts. Deterministic code verifies.