{"slug": "samjho", "title": "SAMJHO", "summary": "A developer built SAMJHO, an open-source, local-first AI document-understanding tool that turns confusing notices, policies and forms into structured summaries, required actions, deadlines and costs. The project runs Gemma 3 1B entirely offline via Ollama, uses PyMuPDF and Tesseract OCR for text and scanned PDFs, and adds an evidence validator that cross-references AI claims against exact document quotes to prevent hallucinations.", "body_md": "*This is a submission for the [Hacktoberfest Weekend Challenge: Build for a Friend](https://dev.to/challenges/hacktoberfest-weekend-2026-10-01)*\n\nI built SAMJHO, a privacy-first, local AI document-understanding MVP designed to simplify confusing notices, policies, and forms into clear, actionable insights for everyday users.\n\nCore Components Built\n\nLocal Document Pipeline (parser.py + OCR): Extracts text page-by-page from standard text PDFs and handles scanned image-based documents using Tesseract OCR.\n\nLocal AI Engine (ai.py + Gemma 3 1B): Powered entirely offline via Ollama to generate structured JSON data—including summaries, required actions, deadlines, and costs—without relying on external proprietary APIs.\n\nEvidence Validator (evidence.py): Powers the PROVE IT feature, cross-referencing AI-generated claims directly against original document quotes to guarantee reliability and prevent hallucinations.\n\nStreamlit Interface: A simple, user-friendly UI featuring transparent processing steps, bilingual support (English and Hindi), and verifiable source inspection.\n\nBy combining local open-weight models with strict evidence validation, SAMJHO bridges the gap between complex paperwork and real-world understanding while keeping all data completely offline.\n\n<!-- Share a deployed link or a video demo. -->[https://drive.google.com/file/d/1euuxA1SqSctYp01NZTVP9wyLy_nPUnxK/view?usp=drive_link](https://drive.google.com/file/d/1euuxA1SqSctYp01NZTVP9wyLy_nPUnxK/view?usp=drive_link)\n\n<!-- Show us the code!  You can embed a GitHub repo directly into your post. --> [https://github.com/ViV1-siNgh/Samjho.git](https://github.com/ViV1-siNgh/Samjho.git)\n\n<!-- Which open-source AI did you use (open-weight models, agent harnesses, frameworks, local inference), and how is your project built around it? --> I built SAMJHO using a strict local-first, modular architecture to ensure complete privacy and reliability.\n\nFirst, I created a robust document parser (parser.py) using PyMuPDF to extract text page-by-page, integrating Tesseract OCR for scanned PDFs. Next, I connected the local open-weight model Gemma 3 1B via Ollama (ai.py) to process the text offline, instructing it to return clean, structured JSON without hallucinating missing details. To guarantee accuracy, I built an evidence validator (evidence.py) that powers a \"PROVE IT\" feature, cross-referencing AI claims against exact document quotes. Finally, I unified the pipeline into a simple Streamlit interface offering bilingual support.\n\n<!-- Why does open innovation matter for what you built? What did it make possible that a closed API wouldn't? -->Open innovation matters because complex, real-world problems—like making bureaucratic paperwork accessible to everyday people—cannot be solved in isolation. By building SAMJHO as an open-source project, open innovation drives several key benefits:\n\nTransparency and Trust: Open-source architectures allow anyone to inspect the codebase, verify that document processing happens entirely locally, and ensure that private data is never sent to external servers.\n\nCommunity Collaboration: It enables developers, designers, and domain experts to build upon existing foundations, share improvements, and adapt solutions to new languages, regions, or document formats.\n\nRapid Iteration: Sharing code and methodologies openly fosters quick feedback loops, turning MVPs into reliable, robust tools much faster than closed-door development.\n\nAccessibility: Open innovation democratizes AI technology, ensuring that practical tools like evidence-backed document explainers remain free, modular, and accessible to everyone who needs them.\n\n<!-- Which partner categories are you entering?  List every one that applies, or remove this section. -->Best Use of Tinker\n\nUse Thinking Machines' Tinker to fine-tune a model for a specific task, and show a clear improvement in performance, latency, or cost over a baseline.\n\nViV1-siNgh - [https://github.com/ViV1-siNgh](https://github.com/ViV1-siNgh)\n\nHimanshu699-cyber - [https://github.com/Himanshu699-cyber](https://github.com/Himanshu699-cyber)", "url": "https://wpnews.pro/news/samjho", "canonical_source": "https://dev.to/vivek_singh_51/samjho-1bpc", "published_at": "2026-10-05 03:36:06+00:00", "updated_at": "2026-10-05 03:42:52.401093+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-tools", "ai-safety", "natural-language-processing"], "entities": ["SAMJHO", "Gemma 3 1B", "Ollama", "Tesseract OCR", "PyMuPDF", "Streamlit", "Hacktoberfest", "ViV1-siNgh"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/samjho", "markdown": "https://wpnews.pro/news/samjho.md", "text": "https://wpnews.pro/news/samjho.txt", "jsonld": "https://wpnews.pro/news/samjho.jsonld"}}