From Comics to Code: Building MyZubster One Piece of Evidence at a Time A developer building the MyZubster evidence platform found that a local RAG pipeline using Qdrant for semantic retrieval and Ollama for inference could retrieve a correct observation yet still generate unsupported claims about it. The team responded by adding a deterministic path that returns exact observations verbatim by ID and replies "Informazione non disponibile nelle fonti MyZubster" when an ID does not exist, rather than substituting a semantically similar record, and exposed the system through an OpenAI-compatible API for use in Open WebUI. From Comics to Code: Building MyZubster One Piece of Evidence at a Time It started with a comic. N4K48 visually represented a journey: a software idea taking shape and imagining its path into the MyZubster metaverse. Those panels were not screenshots of a finished product. They were not proof of a fully decentralized platform. They were a representation of something we wanted to build. Then we started turning that narrative into verifiable software. From Storytelling to Evidence One of the questions that emerged during development was surprisingly simple: How can an AI system distinguish between something that was actually recorded and something that merely sounds plausible? For MyZubster, the answer is increasingly becoming: Start with evidence. An observation enters the system with an identifier. It is persisted, indexed, and made searchable. Qdrant provides semantic retrieval. Ollama provides local AI inference. A RAG layer connects questions, retrieved sources, and generation. But while testing the system, we discovered something more important than simply proving that “RAG works.” We saw how easily a language model can take a correct source and still produce an incorrect answer. For example, MyZubster contains a real observation: ID: 21089771b2a73a9f Description: Test reale MyZubster RC2 - N4K48 Semantic retrieval correctly found that observation. But when we asked the model to describe it, the model could still introduce additional statements that were not supported by the retrieved evidence. That changed the way we approached the problem. Retrieval Is Not the Same as Truth We started separating two problems that are often treated as if they were the same: Finding information and deciding what the system is allowed to claim based on that information. We introduced a deterministic path for observation IDs. If a request contains an existing observation ID, MyZubster prefers the exact observation over semantic similarity. If the ID looks valid but does not exist, the system does not silently search for something similar and use that as a substitute. Instead, it returns: Informazione non disponibile nelle fonti MyZubster. In English: Information is not available in the MyZubster sources. This sounds like a small behavior change. It is not. It means that absence of evidence is no longer treated as permission to generate a plausible substitute. Sometimes the AI Should Not Generate We found the same problem with descriptions. When a user explicitly asks for the description contained in the evidence, why should a language model rewrite it? We tested exactly that. The correct evidence was: Test reale MyZubster RC2 - N4K48 Yet a small local model could turn it into a longer response containing claims that were never present in the source. So we changed the architecture. When MyZubster can answer an authoritative question directly from the retrieved evidence, it can bypass generation and return the evidence verbatim. The result becomes: Test reale MyZubster RC2 - N4K48 Nothing added. Nothing “improved.” Nothing invented. That led us to an important design principle: When the software already knows the fact, we do not necessarily need AI to reinvent the answer. AI becomes useful where interpretation and synthesis are actually required—not as an unavoidable intermediary between the user and every piece of stored evidence. Local AI, Connected to Evidence The current prototype connects several components: User ↓ Open WebUI ↓ OpenAI-compatible MyZubster API ↓ Retrieval / deterministic evidence paths ↓ Qdrant ↓ MyZubster evidence ↓ Ollama when generation is actually needed ↓ Answer We exposed the MyZubster RAG system through an OpenAI-compatible API, allowing Open WebUI to use myzubster-rag as a model. That gave us a practical end-to-end environment in which we could test the entire path from the user interface to the underlying evidence. And the tests uncovered real architectural problems. At one point, an observation existed in persistent storage but was missing from Qdrant. That exposed another important distinction: stored data and searchable data are not automatically the same thing. So we added an explicit observation-index recovery mechanism capable of rebuilding the search index from persisted observations. Again, this is not primarily about making a chatbot answer questions. It is about being able to understand where an answer came from and whether the underlying evidence can be recovered. From Software to Decentralization This is where decentralization enters the story. And this distinction matters: MyZubster does not become decentralized simply because we use local AI, hashes, identifiers, or an internal ledger. A digitally recorded event is not automatically a blockchain event. A database entry is not automatically decentralized evidence. We have deliberately tried to preserve those distinctions in the prototype. The more interesting question is: If evidence originates inside MyZubster today, how could someone verify it tomorrow without having to blindly trust MyZubster itself? That is where decentralization becomes meaningful. Hashes, provenance, identities, attestations, and externally verifiable records can eventually allow evidence to move beyond the trust boundary of a single application. But before decentralizing evidence, we need to know exactly what is being attested. That is one reason we are building the evidence layer first. Decentralization without provenance can simply distribute uncertainty. From the Comic to the Commit History The N4K48 comic represented a vision. The software is slowly turning parts of that vision into properties we can actually test. Some of the latest checkpoints tell that story: 16f0800 fix: add observation index recovery d1aa48f fix: prefer exact observation ID retrieval 8106276 fix: reject unknown observation IDs safely 87a1021 fix: return authoritative descriptions verbatim The interesting part is not the commit count. It is the progression. First, recover the evidence. Then make sure an exact identifier retrieves the exact evidence. Then make sure a nonexistent identifier cannot silently become “something similar.” Then prevent the language model from rewriting an authoritative description when no generation is necessary. Each step removes a little ambiguity from the system. Evidence First, Generation Second This is gradually becoming one of the core ideas behind MyZubster: Observe ↓ Record ↓ Preserve ↓ Retrieve ↓ Verify ↓ Generate only when necessary Most AI systems are designed around the final step. We are increasingly interested in everything that happens before it. Because an eloquent answer is not necessarily a grounded answer. A semantically similar document is not necessarily the requested evidence. A model saying something confidently does not turn it into a recorded fact. And a decentralized record is only useful if we understand what that record actually proves. Where This Is Going The goal of MyZubster is not to build an AI that appears to know everything. The more interesting goal is to build a system that becomes increasingly capable of distinguishing between: what was observed, what was recorded, what can be retrieved, what can be verified, and what the AI generated. There is still a lot to solve. Generic semantic questions can still retrieve partially relevant context. Small local models can still infer things that the evidence does not say. Provenance can become stronger. Verification can extend beyond the boundaries of one application. And decentralization remains a direction to build and test rather than a label to attach prematurely. But that is exactly why this journey is interesting. It started with comic panels imagining a world. Now we are working on the less spectacular—but perhaps more important—part: giving that world a verifiable memory. From comics to evidence. From evidence to software. From software to verification. From verification toward decentralization. And this time, every step should leave evidence behind.