{"slug": "from-comics-to-code-building-myzubster-one-piece-of-evidence-at-a-time", "title": "From Comics to Code: Building MyZubster One Piece of Evidence at a Time", "summary": "A developer building the MyZubster evidence platform found that a local RAG pipeline using Qdrant for semantic retrieval and Ollama for inference could retrieve a correct observation yet still generate unsupported claims about it. The team responded by adding a deterministic path that returns exact observations verbatim by ID and replies \"Informazione non disponibile nelle fonti MyZubster\" when an ID does not exist, rather than substituting a semantically similar record, and exposed the system through an OpenAI-compatible API for use in Open WebUI.", "body_md": "From Comics to Code: Building MyZubster One Piece of Evidence at a Time\n\nIt started with a comic.\n\nN4K48 visually represented a journey: a software idea taking shape and imagining its path into the MyZubster metaverse.\n\nThose panels were not screenshots of a finished product. They were not proof of a fully decentralized platform. They were a representation of something we wanted to build.\n\nThen we started turning that narrative into verifiable software.\n\nFrom Storytelling to Evidence\n\nOne of the questions that emerged during development was surprisingly simple:\n\nHow can an AI system distinguish between something that was actually recorded and something that merely sounds plausible?\n\nFor MyZubster, the answer is increasingly becoming:\n\nStart with evidence.\n\nAn observation enters the system with an identifier. It is persisted, indexed, and made searchable.\n\nQdrant provides semantic retrieval. Ollama provides local AI inference. A RAG layer connects questions, retrieved sources, and generation.\n\nBut while testing the system, we discovered something more important than simply proving that “RAG works.”\n\nWe saw how easily a language model can take a correct source and still produce an incorrect answer.\n\nFor example, MyZubster contains a real observation:\n\nID: 21089771b2a73a9f\n\nDescription: Test reale MyZubster RC2 - N4K48\n\nSemantic retrieval correctly found that observation.\n\nBut when we asked the model to describe it, the model could still introduce additional statements that were not supported by the retrieved evidence.\n\nThat changed the way we approached the problem.\n\nRetrieval Is Not the Same as Truth\n\nWe started separating two problems that are often treated as if they were the same:\n\nFinding information and deciding what the system is allowed to claim based on that information.\n\nWe introduced a deterministic path for observation IDs.\n\nIf a request contains an existing observation ID, MyZubster prefers the exact observation over semantic similarity.\n\nIf the ID looks valid but does not exist, the system does not silently search for something similar and use that as a substitute.\n\nInstead, it returns:\n\nInformazione non disponibile nelle fonti MyZubster.\n\nIn English:\n\nInformation is not available in the MyZubster sources.\n\nThis sounds like a small behavior change.\n\nIt is not.\n\nIt means that absence of evidence is no longer treated as permission to generate a plausible substitute.\n\nSometimes the AI Should Not Generate\n\nWe found the same problem with descriptions.\n\nWhen a user explicitly asks for the description contained in the evidence, why should a language model rewrite it?\n\nWe tested exactly that.\n\nThe correct evidence was:\n\nTest reale MyZubster RC2 - N4K48\n\nYet a small local model could turn it into a longer response containing claims that were never present in the source.\n\nSo we changed the architecture.\n\nWhen MyZubster can answer an authoritative question directly from the retrieved evidence, it can bypass generation and return the evidence verbatim.\n\nThe result becomes:\n\nTest reale MyZubster RC2 - N4K48\n\nNothing added.\n\nNothing “improved.”\n\nNothing invented.\n\nThat led us to an important design principle:\n\nWhen the software already knows the fact, we do not necessarily need AI to reinvent the answer.\n\nAI becomes useful where interpretation and synthesis are actually required—not as an unavoidable intermediary between the user and every piece of stored evidence.\n\nLocal AI, Connected to Evidence\n\nThe current prototype connects several components:\n\nUser\n\n  ↓\n\nOpen WebUI\n\n  ↓\n\nOpenAI-compatible MyZubster API\n\n  ↓\n\nRetrieval / deterministic evidence paths\n\n  ↓\n\nQdrant\n\n  ↓\n\nMyZubster evidence\n\n  ↓\n\nOllama when generation is actually needed\n\n  ↓\n\nAnswer\n\nWe exposed the MyZubster RAG system through an OpenAI-compatible API, allowing Open WebUI to use myzubster-rag as a model.\n\nThat gave us a practical end-to-end environment in which we could test the entire path from the user interface to the underlying evidence.\n\nAnd the tests uncovered real architectural problems.\n\nAt one point, an observation existed in persistent storage but was missing from Qdrant.\n\nThat exposed another important distinction:\n\nstored data and searchable data are not automatically the same thing.\n\nSo we added an explicit observation-index recovery mechanism capable of rebuilding the search index from persisted observations.\n\nAgain, this is not primarily about making a chatbot answer questions.\n\nIt is about being able to understand where an answer came from and whether the underlying evidence can be recovered.\n\nFrom Software to Decentralization\n\nThis is where decentralization enters the story.\n\nAnd this distinction matters:\n\nMyZubster does not become decentralized simply because we use local AI, hashes, identifiers, or an internal ledger.\n\nA digitally recorded event is not automatically a blockchain event.\n\nA database entry is not automatically decentralized evidence.\n\nWe have deliberately tried to preserve those distinctions in the prototype.\n\nThe more interesting question is:\n\nIf evidence originates inside MyZubster today, how could someone verify it tomorrow without having to blindly trust MyZubster itself?\n\nThat is where decentralization becomes meaningful.\n\nHashes, provenance, identities, attestations, and externally verifiable records can eventually allow evidence to move beyond the trust boundary of a single application.\n\nBut before decentralizing evidence, we need to know exactly what is being attested.\n\nThat is one reason we are building the evidence layer first.\n\nDecentralization without provenance can simply distribute uncertainty.\n\nFrom the Comic to the Commit History\n\nThe N4K48 comic represented a vision.\n\nThe software is slowly turning parts of that vision into properties we can actually test.\n\nSome of the latest checkpoints tell that story:\n\n16f0800  fix: add observation index recovery\n\nd1aa48f  fix: prefer exact observation ID retrieval\n\n8106276  fix: reject unknown observation IDs safely\n\n87a1021  fix: return authoritative descriptions verbatim\n\nThe interesting part is not the commit count.\n\nIt is the progression.\n\nFirst, recover the evidence.\n\nThen make sure an exact identifier retrieves the exact evidence.\n\nThen make sure a nonexistent identifier cannot silently become “something similar.”\n\nThen prevent the language model from rewriting an authoritative description when no generation is necessary.\n\nEach step removes a little ambiguity from the system.\n\nEvidence First, Generation Second\n\nThis is gradually becoming one of the core ideas behind MyZubster:\n\nObserve\n\n   ↓\n\nRecord\n\n   ↓\n\nPreserve\n\n   ↓\n\nRetrieve\n\n   ↓\n\nVerify\n\n   ↓\n\nGenerate only when necessary\n\nMost AI systems are designed around the final step.\n\nWe are increasingly interested in everything that happens before it.\n\nBecause an eloquent answer is not necessarily a grounded answer.\n\nA semantically similar document is not necessarily the requested evidence.\n\nA model saying something confidently does not turn it into a recorded fact.\n\nAnd a decentralized record is only useful if we understand what that record actually proves.\n\nWhere This Is Going\n\nThe goal of MyZubster is not to build an AI that appears to know everything.\n\nThe more interesting goal is to build a system that becomes increasingly capable of distinguishing between:\n\nwhat was observed,\n\nwhat was recorded,\n\nwhat can be retrieved,\n\nwhat can be verified,\n\nand what the AI generated.\n\nThere is still a lot to solve.\n\nGeneric semantic questions can still retrieve partially relevant context. Small local models can still infer things that the evidence does not say. Provenance can become stronger. Verification can extend beyond the boundaries of one application. And decentralization remains a direction to build and test rather than a label to attach prematurely.\n\nBut that is exactly why this journey is interesting.\n\nIt started with comic panels imagining a world.\n\nNow we are working on the less spectacular—but perhaps more important—part:\n\ngiving that world a verifiable memory.\n\nFrom comics to evidence.\n\nFrom evidence to software.\n\nFrom software to verification.\n\nFrom verification toward decentralization.\n\nAnd this time, every step should leave evidence behind.", "url": "https://wpnews.pro/news/from-comics-to-code-building-myzubster-one-piece-of-evidence-at-a-time", "canonical_source": "https://dev.to/n4k48/from-comics-to-code-building-myzubster-one-piece-of-evidence-at-a-time-489i", "published_at": "2026-10-03 20:25:32+00:00", "updated_at": "2026-10-03 20:38:17.140703+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "large-language-models", "ai-infrastructure"], "entities": ["MyZubster", "Qdrant", "Ollama", "Open WebUI", "N4K48"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/from-comics-to-code-building-myzubster-one-piece-of-evidence-at-a-time", "markdown": "https://wpnews.pro/news/from-comics-to-code-building-myzubster-one-piece-of-evidence-at-a-time.md", "text": "https://wpnews.pro/news/from-comics-to-code-building-myzubster-one-piece-of-evidence-at-a-time.txt", "jsonld": "https://wpnews.pro/news/from-comics-to-code-building-myzubster-one-piece-of-evidence-at-a-time.jsonld"}}