{"slug": "your-llm-can-t-read-your-documents-our-new-parser-can", "title": "Your LLM Can't Read Your Documents. Our New Parser Can.", "summary": "LM-Kit released LM-Kit Document Parsing, a document parsing engine for LM-Kit.NET and LM-Kit One that scores 86.0 on ParseBench, which the company says is the best result of any self-hostable method. The parser returns typed, positioned elements in reading order — including tables with merged-cell spans, charts as data tables, and bounding boxes with confidence scores — in Markdown, lossless JSON, semantic HTML, and DocLang, and runs entirely on local machines. ParseBench, a public document-parsing benchmark from LlamaIndex, covers more than 2,000 human-verified pages across five dimensions and lists 142 methods on its leaderboard.", "body_md": "Every team that puts an LLM in front of real documents meets the same wall. The demo works on a clean PDF. Then come the annual reports, the scanned contracts, the spec sheets with merged headers, the charts that carry the only copy of a number. The answers get vague, the totals come out wrong, and nobody can say which page a claim came from.\n\nThe model is rarely the problem. The problem is what the model was given.\n\nToday we are releasing [LM-Kit Document Parsing](https://lm-kit.com/solutions/document-intelligence/document-parsing/), the engine that sits between your documents and your LLMs, in [LM-Kit.NET](https://lm-kit.com/products/lm-kit-net/) and [LM-Kit One](https://lm-kit.com/products/lm-kit-one/). It scores 86.0 on ParseBench, the public benchmark for document parsing, the best result of any method you can run yourself. It runs on your machines, and no page ever leaves them.\n\n## Why LLMs struggle with documents\n\nA PDF is not a document. It is a set of drawing instructions: put these glyphs here, fill this rectangle there. Everything a reader understands at a glance, the columns, the table grid, the caption that belongs to a figure, the chart's values, exists only in the picture. Extract the text and all of it is gone.\n\n- **Tables become word soup.** Cells come out in drawing order, so a table's numbers land inside the paragraph next to it, detached from their rows and columns.\n- **Charts disappear.** A bar chart contributes its axis ticks and nothing else. The values it plots never reach the model.\n- **Reading order breaks.** Two columns are read across, line by line, interleaving two arguments into one.\n- **The text layer lies.** Fonts that encode the wrong characters, words hidden under images, the guesses of another engine's OCR on a scanned page: all of it is text, and none of it is what the page says.\n\nFeed that into retrieval and the errors are already baked in before a single embedding is computed. That is why so many RAG projects stall: [most RAG errors start before retrieval](https://lm-kit.com/solutions/document-intelligence/document-parsing/#rag), at the moment a page is flattened into text.\n\n            LLMs don't read pages. They read what you give them. Document understanding is the layer they lack.\n        \n\n## What the parser returns\n\nThe parser reads a page the way a person does and returns what it found as data: every element typed and positioned, in reading order.\n\n- **Structure.** Titles and heading levels that stay consistent across pages, paragraphs, lists, captions bound to their figures, footnotes and running heads kept apart from the body.\n- **Tables, whole.** Merged cells survive as row and column spans, negatives keep their parentheses, totals stay bold, and a table split across a layout break is joined back.\n- **Charts, as data.** Bars, lines, pies, stacks and multi-panel figures come back as the table they plot, with their legend and units.\n- **Grounding.** Every element carries its bounding box, its place in the reading order and a confidence, so any answer built on it can cite the exact region it came from.\n\nOne parse, every standard format\n\n- Markdown for LLM prompts and RAG chunks, with tables as HTML and formulas as LaTeX.\n- Lossless JSON with a published schema, for pipelines and storage in any language.\n- Semantic HTML for display and review, and DocLang, the open document markup of the LF AI & Data Foundation.\n\nThe fastest way to see it is to use it. The [Parse Studio](https://lm-kit.com/solutions/document-intelligence/document-parsing/#studio) on the product page shows the parser's real output on three documents: tap a region, play the reading order, read the Markdown and JSON your code receives.\n\n## The results\n\n[ParseBench](https://www.parsebench.ai/) is a public document-parsing benchmark from LlamaIndex: more than 2,000 human-verified pages and five dimensions that matter to an AI system reading documents, namely tables, charts, content faithfulness, semantic formatting and visual grounding. Its leaderboard puts 142 methods side by side: cloud parsing services, the hyperscalers' document APIs, frontier models and open-weight parsers.\n\n**#1** of every method you can run yourself, at all three effort levels\n\n**#1** on visual grounding, out of all 142 methods on the board\n\n**#3** overall, behind only a cloud service billed per page\n\n### First of everything you can run yourself\n\nThe strongest self-hosted method on the board scores 78.3 overall. LM-Kit scores 86.0 at High, 84.4 at Medium and 80.1 at Low: even the fastest level is ahead of every other parser you can host on your own hardware.\n\n### First on visual grounding, of all 142 methods\n\nVisual grounding asks whether each element comes back in the right place on the page. It decides whether an answer can point to its exact source, which is what makes a RAG answer or an extracted value checkable. The best score on the public board is 84.3. LM-Kit reaches 87.5 at High and 86.0 at Medium, ahead of every method listed, cloud services included.\n\n### Against the platforms you are likely evaluating\n\nThe document APIs of the major clouds score 59.6 (Azure Document Intelligence, Layout), 50.4 (Google Cloud Document AI) and 47.9 (AWS Textract). Databricks AI Parse scores 60.7, Mistral OCR 68.2, Reducto 73.0. The strongest frontier models, run at their highest reasoning settings, land between 75.0 and 79.8: Gemini 3 Flash, GPT-5.6 Sol and Claude Opus 5.5. LM-Kit High is ahead of all of them by more than six points, without sending a page anywhere.\n\nThe mean of the five ParseBench dimensions: tables, charts, content faithfulness, semantic formatting and visual grounding.\n\nWhether each element comes back in the right place, so an answer can point to its exact source. No method scores higher.\n\nOnly the methods you can run on your own hardware. All three LM-Kit levels lead them.\n\n    Ranks among all 142 methods on the public board plus LM-Kit's three levels; a dotted row marks methods not shown.\n    LM-Kit scores: the public ParseBench evaluation code and dataset, charts graded without the optional LLM judge.\n    Other methods: [parsebench.ai leaderboard](https://www.parsebench.ai/#leaderboard), October 3, 2026.\n\nThe two methods ranked above LM-Kit High are the agentic tiers of LlamaParse, a cloud service billed per page. LM-Kit outscores both on visual grounding. The public [LM-Kit One](https://lm-kit.com/products/lm-kit-one/) container reproduces the High score with its default configuration.\n\n## Not one model: an engine that checks every reading\n\nThe parser is not a single large model. It orchestrates OCR, vision-language readers, layout detection, natural language processing and the page's own geometry, and weighs every reading with scoring equations and metrics of our own before one stands.\n\n- **Route before reading.** Each page and region is sorted before any model runs. Text a trusted text layer prints is served from it; tables, figures, formulas and scans go to the readers.\n- **Read, then read again.** A vision reader transcribes the page, and a second, independent reader takes figures, tables and doubtful text again.\n- **Settle by evidence.** Chart values are checked against the drawing, labels against the words the figure prints, tables against the grid the page draws. A reading the page contradicts does not stand.\n- **Every token through Dynamic Sampling.** At inference time,[Dynamic Sampling](https://lm-kit.com/solutions/local-inference/dynamic-sampling) tracks the structure being written, from the taxonomy of elements to a table's grid and a chart's series, and steers each token to fit it.\n\nThe same discipline handles what real-world PDFs do to a parser. A text layer is treated as a claim, not a fact: fonts that encode the wrong characters are caught and decoded, text drawn where nobody can see it is dropped, the hidden layer of a searchable scan is a hint rather than the truth, and bold is measured from the rendered stroke. Pages photographed at an angle are set straight before they are read, and sideways pages are turned upright. Untagged or badly tagged files parse like any other, because the structure comes from the page itself.\n\nThis is how small models out-read much larger ones, and it is why the engine keeps getting better. Document AI research moves every month. We keep integrating the strongest new readers, layout models and techniques, and a component ships only when it improves the whole. Your code does not change: the same three settings get faster and more accurate from one release to the next.\n\n## One setting for speed\n\nA parse takes one decision: an effort level. Every other choice belongs to the engine.\n\n**44.4** pages a minute at Low, for high-volume ingestion\n\n**29.6** pages a minute at Medium, the default, at 98% of High's score\n\n**19.8** pages a minute at High, for filings and reports\n\nThose rates come from the full ParseBench dataset on a single desktop GPU, four documents at a time. The parser also runs on CPU, and on CUDA, Vulkan and Metal, on Windows, Linux and macOS.\n\n## Built for RAG, extraction and agents\n\nParsing is rarely the goal. It is what makes the next step trustworthy.\n\n- **RAG that holds up.** Chunks follow headings and keep a table or a chart whole, table cells and chart values become text an index can match, and every chunk keeps its page and box for citations.[Document RAG](https://lm-kit.com/solutions/document-intelligence/document-rag/) and the[parsed-document RAG demo](https://github.com/LM-Kit/lm-kit-net-samples/tree/main/console_net/document-intelligence/document-parsing/parsed_document_rag) show the pattern.\n- **Extraction you can review.**[Structured data extraction](https://lm-kit.com/solutions/document-intelligence/structured-data-extraction/) turns documents into schema-valid fields with a confidence on each, so a reviewer only looks where it is needed.\n- **Agents that can read.** LM-Kit One exposes the parser as an MCP tool, so an agent can ingest a file and call`document_parse` under the permissions your administrators set.\n\n## Edge, local, air-gapped\n\nCloud parsers bill every page and see every page. At their listed prices, the two methods that score above LM-Kit High would cost $13,000 and $56,000 to parse a million pages. LM-Kit has no per-page API charge, and the documents never leave your network.\n\n- **Edge.** Embed the parser in a .NET application on laptops, workstations, scanning stations and factory PCs, and parse where documents are captured.\n- **Local.** Serve it from LM-Kit One on your private servers, to any language over REST, from the command line, and to agents over MCP.\n- **Air-gapped.** Ship the single model file with your application. No call home, nothing to open in the firewall, nothing to download at run time.\n\nFor organizations handling contracts, medical records or financial filings, that is the difference between a project legal can approve and one it cannot. We set out what private means, and why local alone is not enough, in [Private AI or local AI?](https://lm-kit.com/blog/private-ai-or-local-ai/)\n\n## Two products, one parser\n\nIn LM-Kit.NET, a parse is a few lines of C#:\n\n``` js\nusing LMKit.Document.Parsing;\n\nusing var parser = new DocumentParser(ParsingEffort.Medium);\nParsedDocument document = parser.Parse(\"annual-report.pdf\");\n\nFile.WriteAllText(\"report.md\", document.ToMarkdown());\nFile.WriteAllText(\"report.json\", document.ToJson());\n```\n\nIn LM-Kit One, the same parser answers at `POST /lmkit/v1/document-parsing`, runs over a whole folder with `lmkit run document-parsing`, and is offered to agents as an MCP tool. The LM-Kit One Playground ships a Parse workbench that shows every element on the page and in the output, side by side. The [product page](https://lm-kit.com/solutions/document-intelligence/document-parsing/#code) has an example for each.\n\n## What comes next\n\nThe team behind this engine brings more than 25 years of PDF, imaging and document-analysis experience, and its leadership has built category-leading document processing products before. Document Parsing is the work we care about most, and it will keep improving with every release: faster on the same hardware, more faithful on the documents that matter, with no change to the code you write today.\n\nBring your hardest documents. That is what we built it for.\n\n### Parse your own documents\n\nFree to build and evaluate, on your hardware, with nothing leaving it.\n\nDocuments your current pipeline gets wrong? [Tell us about them](https://lm-kit.com/contact/#contact-form): what they are, what you need out of them, and what a useful result would look like.", "url": "https://wpnews.pro/news/your-llm-can-t-read-your-documents-our-new-parser-can", "canonical_source": "https://lm-kit.com/blog/introducing-document-parsing/", "published_at": "2026-10-03 00:00:00+00:00", "updated_at": "2026-10-03 09:38:48.648215+00:00", "lang": "en", "topics": ["ai-tools", "large-language-models", "ai-products", "structured-data"], "entities": ["LM-Kit", "LM-Kit Document Parsing", "LM-Kit.NET", "LM-Kit One", "ParseBench", "LlamaIndex", "LF AI & Data Foundation", "DocLang"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/your-llm-can-t-read-your-documents-our-new-parser-can", "markdown": "https://wpnews.pro/news/your-llm-can-t-read-your-documents-our-new-parser-can.md", "text": "https://wpnews.pro/news/your-llm-can-t-read-your-documents-our-new-parser-can.txt", "jsonld": "https://wpnews.pro/news/your-llm-can-t-read-your-documents-our-new-parser-can.jsonld"}}