{"slug": "parsing-korea-s-hwp-government-documents-into-markdown-for-llms", "title": "Parsing Korea's HWP government documents into Markdown for LLMs", "summary": "A civil servant at Seoul's Gwangjin-gu District Office built kordoc, an MIT-licensed TypeScript library, CLI and MCP server that converts Korean HWP/HWPX government documents into Markdown and structured data for LLM pipelines. The tool parses HWP 3.x/5.x, HWPX, HWPML, PDF, Office and image formats without Hancom Office installed, and exposes 17 tools to agents over MCP. On its 4.21.14 release verification, all 10,342 visible tables from 2,424 HWPX documents matched in structure, and it scored 0.963 overall on the external opendataloader-bench of 200 PDFs.", "body_md": "I am a local civil servant at the Gwangjin-gu District Office in Seoul. I spent seven years wrestling HWP files at work, and when I started wiring LLMs into that work, the files were the first wall I hit. Almost everything I read and write on the job is an HWP or HWPX document. RAG pipelines and AI agents want Markdown. So I built [kordoc](https://github.com/chrisryugj/kordoc), an MIT-licensed TypeScript library, CLI and MCP server that converts Korean documents into Markdown and structured data.\n\nHWP is the Hancom word processor format, and there is more than one of it:\n\n`head.xml` for styles and numbering. When the ZIP central directory is broken, kordoc scans the local file headers directly.\nNone of this needs Hancom Office installed. kordoc reads the files directly.\n\nIt converts HWP 3.x/5.x, HWPX, HWPML, PDF, XLS/XLSX, DOCX, PPTX and PNG/JPG/WebP into Markdown and structured data. Beyond reading, it does block and cell diffs, format-preserving HWPX/HWP patches, Markdown to HWPX generation with government-document presets, form filling, previews and PII masking. Through MCP it exposes 17 tools to an agent.\n\nNode.js 20+ on macOS, Linux or Windows.\n\n```\nnpm install kordoc\n\nnpx kordoc document.hwpx -o document.md\nnpx kordoc *.pdf --jobs 4 -d ./output\nnpx kordoc document.pdf --format json --pages 1-3\nnpx kordoc scan.pdf --ocr -o scan.md\n```\n\nAs a library:\n\n``` js\nimport { parse } from \"kordoc\"\n\nconst result = await parse(\"document.hwpx\")\nif (result.success) {\n  console.log(result.markdown)\n  // result.blocks: structured data; result.metadata: document metadata\n}\n```\n\nTo give an agent access, one command registers the MCP server with installed clients such as Claude Desktop, Claude Code, Cursor and Codex:\n\n```\nnpx -y kordoc setup\n```\n\nIn Claude Code you can also install it as a plugin:\n\n```\n/plugin marketplace add chrisryugj/kordoc\n/plugin install kordoc@kordoc\n```\n\nThere is now a Python SDK on PyPI. It does not bundle the engine. It starts a resident kordoc engine worker and reuses it, so you need Node.js 20+ and the npm package on the machine, plus Python 3.11+:\n\n```\nnpm i -g kordoc        # engine\npip install kordoc     # SDK\npython\nfrom kordoc import KordocClient\n\nwith KordocClient() as client:\n    result = client.parse(\"document.hwpx\", table_format=\"gfm\")\n    if result.success:\n        print(result.markdown)\n```\n\nThe documents I handle are full of tables with merged and nested cells. kordoc lays out tables in two passes: first it computes the grid size from colSpan/rowSpan, then it places cells. For RAG indexing, `--table-format gfm` (API: `tableFormat: \"gfm\"`) emits merged and nested tables as GFM pipe tables with no HTML, moving nested tables out with parent/child markers. `--format chunks` gives structure chunks with heading breadcrumbs plus standalone table chunks.\n\nOn the 4.21.14 release verification (2026-10-10), all 10,342 visible tables from 2,424 original HWPX documents matched in structure (100%). On 1,130 HWP 5.x and HWPX pairs with 4,315 tables, paired-document table structure match was also 100%.\n\nNot every document arrives as the HWP original; some arrive only as a PDF export. I score kordoc's PDF output against the original HWPX as ground truth. On 744 pairs: character recall 99.84%, precision 99.64%, reading order 99.16%, word-boundary F1 98.89%. On PDF tables (708 pairs, 2,331 tables): detection 99.83%, structure match 97.98%, cell F1 0.991.\n\nOn the external opendataloader-bench (200 PDFs), the 4.21.14 release scores 0.963 overall with defaults and 0.939 with OCR off. In the 2026-09-29 comparison against 12 published parsers it ranked first. The full methodology, exclusions and per-option results are in the [benchmark doc](https://github.com/chrisryugj/kordoc/blob/main/docs/benchmarks-en.md).\n\nkordoc is meant to run on air-gapped networks too. `KORDOC_OFFLINE=1` blocks all outbound traffic, and OCR runs on the local CPU with no API key. For MCP, `KORDOC_ROOT` restricts which files the server can touch. On 53 documents and 102 pages, OCR measures CER 0.0412 and character recall 99.02%.\n\nTable structure, cell content and visual fidelity are separate metrics; a 100% structure score does not mean every cell's text or every page's look is perfect. PDF tables are at 97.98% structure match, not 100%. Default OCR only runs when the OCR model is already cached; otherwise nothing is downloaded and you get `NEEDS_OCR` / `SKIPPED_IMAGE` warnings. Format-preserving patches skip edits they cannot apply and report the reason. The HWPX cell-edit path requires the same number of nonempty lines, so blank lines, added or deleted lines, literal `<br>` and ambiguous mappings stay unsupported.\n\nThe second tool I lean on at work is [korean-law-mcp](https://github.com/chrisryugj/korean-law-mcp), an MCP server and CLI for Korea's official legal database (the 법제처 Open API). It wraps 42 APIs into 10 tools covering statutes, precedents, administrative rules, local ordinances, treaties and legal interpretations. The part most relevant to LLM work is citation verification: `legal_analysis(mode=\"verify_citations\")` checks the statute articles and case numbers cited in a text against the official source, for existence and for content. It uses kordoc to parse statute annexes, and it is listed in the official MCP Registry. The API key it needs is free.\n\n```\nnpx --ignore-scripts --omit=optional korean-law-mcp setup\n```\n\nIf you work with Korean documents, the most useful thing you can send me is a real file that breaks, as an issue. Stars help other people find it too: [github.com/chrisryugj/kordoc](https://github.com/chrisryugj/kordoc).", "url": "https://wpnews.pro/news/parsing-korea-s-hwp-government-documents-into-markdown-for-llms", "canonical_source": "https://dev.to/chrisryugj/parsing-koreas-hwp-government-documents-into-markdown-for-llms-1g8k", "published_at": "2026-10-10 13:11:35+00:00", "updated_at": "2026-10-10 13:15:56.174040+00:00", "lang": "en", "topics": ["ai-tools", "agent-protocols", "structured-data", "natural-language-processing", "developer-tools"], "entities": ["kordoc", "Gwangjin-gu District Office", "Hancom", "Claude Desktop", "Claude Code", "Cursor", "Codex", "PyPI"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/parsing-korea-s-hwp-government-documents-into-markdown-for-llms", "markdown": "https://wpnews.pro/news/parsing-korea-s-hwp-government-documents-into-markdown-for-llms.md", "text": "https://wpnews.pro/news/parsing-korea-s-hwp-government-documents-into-markdown-for-llms.txt", "jsonld": "https://wpnews.pro/news/parsing-korea-s-hwp-government-documents-into-markdown-for-llms.jsonld"}}