One Interface, Seven Formats: How Solon AI Turns Files, Web Pages, and Even Database Schemas into RAG Documents A developer released solon-ai-rag-loaders, a Java library that implements a single three-method DocumentLoader interface across seven Maven sub-modules for Markdown, PDF, Word, Excel, HTML, PowerPoint, and database DDL sources. Each loader picks a format-specific default chunking unit — sections for Markdown, pages for PDF, paragraphs for Word, sheets batched at 200 rows for Excel, and one SHOW CREATE TABLE per table for DDL — rather than exposing a single global chunk-size setting. The Markdown loader parses documents into a commonmark AST and attaches heading, code-block, and blockquote metadata to each resulting Document. Every RAG pipeline starts the same way: you have stuff, and the model needs Document s. The interesting question is how far that idea stretches. Solon AI answers it with a deliberately small contract — and then pushes it across seven formats, including one you probably haven't tried feeding to a retriever: your database schema. This is a source-code tour of solon-ai-rag-loaders . All claims below are checked against the current source tree; where a class behaves in a way you wouldn't guess from its name, I'll point it out. public interface DocumentLoader { DocumentLoader additionalMetadata String key, Object value ; DocumentLoader additionalMetadata Map