{"slug": "one-interface-seven-formats-how-solon-ai-turns-files-web-pages-and-even-database", "title": "One Interface, Seven Formats: How Solon AI Turns Files, Web Pages, and Even Database Schemas into RAG Documents", "summary": "A developer released solon-ai-rag-loaders, a Java library that implements a single three-method DocumentLoader interface across seven Maven sub-modules for Markdown, PDF, Word, Excel, HTML, PowerPoint, and database DDL sources. Each loader picks a format-specific default chunking unit — sections for Markdown, pages for PDF, paragraphs for Word, sheets batched at 200 rows for Excel, and one SHOW CREATE TABLE per table for DDL — rather than exposing a single global chunk-size setting. The Markdown loader parses documents into a commonmark AST and attaches heading, code-block, and blockquote metadata to each resulting Document.", "body_md": "Every RAG pipeline starts the same way: you have stuff, and the model needs `Document` s.\n\nThe interesting question is how far that idea stretches. Solon AI answers it with a deliberately small contract — and then pushes it across seven formats, including one you probably haven't tried feeding to a retriever: your database schema.\n\nThis is a source-code tour of `solon-ai-rag-loaders`. All claims below are checked against the current source tree; where a class behaves in a way you wouldn't guess from its name, I'll point it out.\n\n```\npublic interface DocumentLoader {\n    DocumentLoader additionalMetadata(String key, Object value);\n    DocumentLoader additionalMetadata(Map<String, Object> metadata);\n    List<Document> load() throws IOException;\n}\n```\n\nThat's the entire contract: two metadata methods and one `load()`. No provider field, no API key, no vendor. Seven Maven sub-modules implement it (`solon-ai-load-markdown`, `-pdf`, `-word`, `-excel`, `-html`, `-ppt`, `-ddl`), each pulling only its own parsing dependency — commonmark, PDFBox, POI, jsoup, Tika.\n\nThe base class `AbstractOptionsDocumentLoader` adds the options pattern with two entry points:\n\n``` php\nMarkdownLoader loader = new MarkdownLoader(file)\n        .options(o -> o.codeBlockAsNew(true));\n// or, if you already hold an Options instance:\nloader.options(myOptions);\n```\n\nA `SupplierEx<InputStream>` constructor appears in every loader, so your source can be a file, a URL, a byte array, or anything else that can produce a stream lazily.\n\nSeven loaders, and no single \"chunk size\" knob. Instead, each loader picks its default unit of meaning — and the defaults disagree on purpose:\n\n| Loader | Default unit | Default mode | \n|---|---|---|\n| `MarkdownLoader` | Section (per heading) | AST walk, headings always split | \n| `PdfLoader` | Page | `LoadMode.PAGE` | \n| `WordLoader` | Paragraph | `LoadMode.PARAGRAPH` | \n| `PptLoader` | Whole document | `LoadMode.SINGLE` | \n| `ExcelLoader` | Sheet, batched at 200 rows | JSON rows | \n| `HtmlSimpleLoader` | Whole page | Single document | \n| `DdlLoader` | Table | One `SHOW CREATE TABLE` each | \n\nThat asymmetry is the design. A paragraph is the natural retrieval unit for prose; a page is the natural unit for a PDF; a slide deck usually makes more sense as one document; a table is a complete thought. You *can* override the defaults (`PdfLoader` goes `SINGLE`, `WordLoader` goes `SINGLE`, `PptLoader` splits on `\"\\n\\n\\n\"`), but the out-of-the-box behavior already encodes a per-format answer to \"what is a chunk here?\"\n\n`MarkdownLoader` doesn't slice text with regexes. It parses the document with commonmark into an AST and walks it with a visitor:\n\n`horizontalLineAsNew`, `blockquoteAsNew`, `codeBlockAsNew`.` codeBlockAsNew(true)`, the code block starts its own document. Either way, a fenced block `category=header_1..6` with a `title`, `category=code_block` with `lang`, or `category=blockquote`.\nOne nuance worth knowing before you rely on metadata: the visitor writes `title`/` category` onto the *current* document while walking. If a section has no heading text before its content, the metadata simply won't be there for that chunk. Fine for retrieval; worth remembering if you build UI on top of it.\n\n`PdfLoader` (PDFBox) defaults to one `Document` per page, each stamped with `page`, `total_pages`, and a `summary` of `\"Page 3\"` — handy in a search UI. Switch to `LoadMode.SINGLE` and you get the whole file as one document, pages joined by `\"\\n\\f\"`, with just a `pages` count.\n\n`WordLoader` handles both binary eras: it checks the stream with POI's `FileMagic` and routes `.docx` (OOXML) and legacy `.doc` (OLE2) to different readers. It defaults to paragraph mode — one `Document` per paragraph — with a `SINGLE` escape hatch.\n\n`ExcelLoader` (POI + snack4) treats the first non-empty row of a sheet as the header row, then maps every following row to `{column: value}` and serializes batches as JSON documents. Two defaults shape its behavior:\n\n`documentMaxRows(-1)` to keep one document per sheet.`break` s, so anything after the first blank row is silently ignored — by design, trailing blank rows shouldn't kill the parse, but data below a blank row won't be indexed. Keep that in mind with hand-edited spreadsheets.\nFormula cells are read as their formula text, not computed values.\n\n`PptLoader` doesn't parse slide XML itself. It hands the stream to Apache Tika's `AutoDetectParser` and gets body text back. Default is `SINGLE` — the whole deck as one document; `PAGE` mode splits on `\"\\n\\n\\n\"` if your decks have predictable slide breaks.\n\nThis is the one that changes how you think about the pipeline. `DdlLoader` connects to a plain `DataSource` (no ORM, no entities) and emits **one `Document` per table** containing its DDL:\n\n```\nDdlLoader loader = new DdlLoader(dataSource);   // MySQL config built in\nloader.options(o -> o.loadOptions(\"shop\", null)); // schema only: all its tables\nList<Document> docs = loader.load();\n```\n\nThree granularities via `loadOptions(schema, table)`: whole instance, one schema, one table. The default configuration is MySQL (`information_schema` + `SHOW CREATE TABLE`, system schemas excluded), but every SQL string is a template — the loader runs them through Solon's own expression engine (`SnEL.evalTmpl`), so you can rewire it for another database by replacing the template set in a `DdlLoadConfig`.\n\nOne detail I like: `SHOW CREATE TABLE` returns `CREATE TABLE \\` order`(...)` — a table name that's only meaningful inside its schema. The loader **rewrites the header** to `CREATE TABLE \\` shop`.\\` order`(...)` so every retrieved DDL document is self-describing, and stamps `metadata(\"table\", \"order\")` so your filter layer can target tables directly.\n\nThe use case writes itself: point it at production (read-only!), and your AI assistant retrieves schema facts instead of hallucinating column names.\n\n`load()`: One Shape Downstream\nWhatever the format, `load()` hands you `List<Document>` — content plus metadata plus the fluent fields (`title`, `url`, `summary`, `id`, `embedding`, `score`). From here, everything is format-agnostic: embed, store in a `Repository`, attach as a tool. The loaders are the only place in the pipeline where format-specific knowledge lives.\n\nTo be fair to your architecture review: if your corpus is already clean Markdown, you might not need seven loaders — Solon AI's splitter story covers embedding-time splitting separately. The loaders earn their keep when sources are heterogeneous (office files, web pages, live schema) or when the natural unit (page, paragraph, table) should decide the chunk, not a character count.\n\n`solon-ai-rag-loaders` is a good example of a small contract held firmly: three methods, seven implementations, and per-format defaults that encode real opinions instead of one generic knob. The DDL loader alone is worth a look if you build assistants that need to talk about your database accurately.\n\n`solon-ai-rag-loaders`.", "url": "https://wpnews.pro/news/one-interface-seven-formats-how-solon-ai-turns-files-web-pages-and-even-database", "canonical_source": "https://dev.to/solonjava/one-interface-seven-formats-how-solon-ai-turns-files-web-pages-and-even-database-schemas-into-e0n", "published_at": "2026-10-10 02:14:50+00:00", "updated_at": "2026-10-10 02:29:13.288036+00:00", "lang": "en", "topics": ["ai-tools", "developer-tools", "large-language-models", "ai-infrastructure"], "entities": ["Solon AI", "solon-ai-rag-loaders", "commonmark", "PDFBox", "Apache POI", "jsoup", "Apache Tika", "snack4"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/one-interface-seven-formats-how-solon-ai-turns-files-web-pages-and-even-database", "markdown": "https://wpnews.pro/news/one-interface-seven-formats-how-solon-ai-turns-files-web-pages-and-even-database.md", "text": "https://wpnews.pro/news/one-interface-seven-formats-how-solon-ai-turns-files-web-pages-and-even-database.txt", "jsonld": "https://wpnews.pro/news/one-interface-seven-formats-how-solon-ai-turns-files-web-pages-and-even-database.jsonld"}}