SAMJHO
A developer built SAMJHO, an open-source, local-first AI document-understanding tool that turns confusing notices, policies and forms into structured summaries, required actions, deadlines and costs. …
A developer built SAMJHO, an open-source, local-first AI document-understanding tool that turns confusing notices, policies and forms into structured summaries, required actions, deadlines and costs. …
A developer built Study Saathi, a local AI study companion that converts uploaded PDF and TXT notes into summaries, practice MCQs, and flashcards. The app runs entirely on-device via Ollama with the q…
A developer built RoastMyResume, an open-source AI tool that delivers brutally honest, section-by-section resume feedback and job-description matching to help a friend who kept getting rejected from i…
A developer built NoteMate AI, an open-source, local-first note assistant that transcribes audio and extracts PDF text entirely on-device using Gemma 3 via Ollama and faster-whisper's CTranslate2 impl…
Jaival Suthar released Knowledge Vault (M1), a local-first PDF retrieval system built on PyMuPDF text extraction, token-aware recursive chunking, BAAI/bge-small-en-v1.5 embeddings stored in Qdrant, an…
A developer built Sahayak, a RAG-based legal assistant that answers tenancy, consumer-rights and contract questions and summarizes uploaded PDFs in plain language, grounding every response in retrieve…
Developer luckmanqasim built PlanLint, an open source AEC compliance checker that restricts large language models to extracting structure and labels while a pure-Python checker serves as the sole verd…
A developer built a Windows desktop application that detects invisible text hidden in PDFs — including white-on-white text, sub-point font sizes, and PDF text render mode 3 — by rendering each page an…
A developer using LlamaIndex reports that embedding models such as BAAI/bge-small-en-v1.5 fail on scientific papers containing special characters, throwing a TypeError: TextEncodeInput must be Union[T…
A developer spent two months building Vestibule, an open-source Python framework for the boring layer of RAG ingestion, with four AI agents handling design, review, implementation, and code review thr…
A developer's integration with Claude Code and markitdown falsely reported a PDF as empty because exit codes only indicate process completion, not content yield, and byte thresholds miss near-misses l…
Google's Gemini 2.5 Flash and text-embedding-004 models, combined with ChromaDB and PyMuPDF, form a RAG pipeline that processes approximately 2GB of data (500,000 text chunks) for technical documentat…
A developer using Microsoft's markitdown converter discovered that a PDF containing only raster images produced a zero-byte output with exit code 0, leading an AI model to incorrectly report the docum…
A developer built a custom AI PDF reader in Python, starting with a Jupyter prototype and evolving it into tested modules. The reader uses PyMuPDF for rendering and search, with features like bookmark…
A developer has converted a Python PDF reader prototype into a desktop application using PySide6, adding features such as page rendering, search, bookmarks, notes, and annotations. The application sep…
A developer built a custom multimodal Retrieval-Augmented Generation (RAG) system that reads and understands thousands of PDFs using open-source AI, extracting text and images, chunking content, and e…
A RAG pipeline using LlamaIndex and the BGE-small-en-v1.5 embedding model crashes with a 'TextEncodeInput must be Union[TextInputSequence, Tuple[InputSequence, InputSequence]]' error when processing s…
A developer improved resume parsing accuracy from 65% to 94% by shifting date arithmetic from an LLM prompt to a Python post-processing script, using a two-step chain that first extracts raw date rang…
Baidu's Unlimited-OCR, a 3B-parameter vision-language model, enables end-to-end OCR for high-resolution images and multi-page PDFs without separate layout analysis, according to a tutorial published o…
A new open-source CLI tool, PDF Batch Translator, translates English PDFs into Japanese Markdown while stripping non-body content and preserving structure. Developed by an unnamed creator, the tool us…