# Parsing Korea's HWP government documents into Markdown for LLMs

> Source: <https://dev.to/chrisryugj/parsing-koreas-hwp-government-documents-into-markdown-for-llms-1g8k>
> Published: 2026-10-10 13:11:35+00:00

I am a local civil servant at the Gwangjin-gu District Office in Seoul. I spent seven years wrestling HWP files at work, and when I started wiring LLMs into that work, the files were the first wall I hit. Almost everything I read and write on the job is an HWP or HWPX document. RAG pipelines and AI agents want Markdown. So I built [kordoc](https://github.com/chrisryugj/kordoc), an MIT-licensed TypeScript library, CLI and MCP server that converts Korean documents into Markdown and structured data.

HWP is the Hancom word processor format, and there is more than one of it:

`head.xml` for styles and numbering. When the ZIP central directory is broken, kordoc scans the local file headers directly.
None of this needs Hancom Office installed. kordoc reads the files directly.

It converts HWP 3.x/5.x, HWPX, HWPML, PDF, XLS/XLSX, DOCX, PPTX and PNG/JPG/WebP into Markdown and structured data. Beyond reading, it does block and cell diffs, format-preserving HWPX/HWP patches, Markdown to HWPX generation with government-document presets, form filling, previews and PII masking. Through MCP it exposes 17 tools to an agent.

Node.js 20+ on macOS, Linux or Windows.

```
npm install kordoc

npx kordoc document.hwpx -o document.md
npx kordoc *.pdf --jobs 4 -d ./output
npx kordoc document.pdf --format json --pages 1-3
npx kordoc scan.pdf --ocr -o scan.md
```

As a library:

``` js
import { parse } from "kordoc"

const result = await parse("document.hwpx")
if (result.success) {
  console.log(result.markdown)
  // result.blocks: structured data; result.metadata: document metadata
}
```

To give an agent access, one command registers the MCP server with installed clients such as Claude Desktop, Claude Code, Cursor and Codex:

```
npx -y kordoc setup
```

In Claude Code you can also install it as a plugin:

```
/plugin marketplace add chrisryugj/kordoc
/plugin install kordoc@kordoc
```

There is now a Python SDK on PyPI. It does not bundle the engine. It starts a resident kordoc engine worker and reuses it, so you need Node.js 20+ and the npm package on the machine, plus Python 3.11+:

```
npm i -g kordoc        # engine
pip install kordoc     # SDK
python
from kordoc import KordocClient

with KordocClient() as client:
    result = client.parse("document.hwpx", table_format="gfm")
    if result.success:
        print(result.markdown)
```

The documents I handle are full of tables with merged and nested cells. kordoc lays out tables in two passes: first it computes the grid size from colSpan/rowSpan, then it places cells. For RAG indexing, `--table-format gfm` (API: `tableFormat: "gfm"`) emits merged and nested tables as GFM pipe tables with no HTML, moving nested tables out with parent/child markers. `--format chunks` gives structure chunks with heading breadcrumbs plus standalone table chunks.

On the 4.21.14 release verification (2026-10-10), all 10,342 visible tables from 2,424 original HWPX documents matched in structure (100%). On 1,130 HWP 5.x and HWPX pairs with 4,315 tables, paired-document table structure match was also 100%.

Not every document arrives as the HWP original; some arrive only as a PDF export. I score kordoc's PDF output against the original HWPX as ground truth. On 744 pairs: character recall 99.84%, precision 99.64%, reading order 99.16%, word-boundary F1 98.89%. On PDF tables (708 pairs, 2,331 tables): detection 99.83%, structure match 97.98%, cell F1 0.991.

On the external opendataloader-bench (200 PDFs), the 4.21.14 release scores 0.963 overall with defaults and 0.939 with OCR off. In the 2026-09-29 comparison against 12 published parsers it ranked first. The full methodology, exclusions and per-option results are in the [benchmark doc](https://github.com/chrisryugj/kordoc/blob/main/docs/benchmarks-en.md).

kordoc is meant to run on air-gapped networks too. `KORDOC_OFFLINE=1` blocks all outbound traffic, and OCR runs on the local CPU with no API key. For MCP, `KORDOC_ROOT` restricts which files the server can touch. On 53 documents and 102 pages, OCR measures CER 0.0412 and character recall 99.02%.

Table structure, cell content and visual fidelity are separate metrics; a 100% structure score does not mean every cell's text or every page's look is perfect. PDF tables are at 97.98% structure match, not 100%. Default OCR only runs when the OCR model is already cached; otherwise nothing is downloaded and you get `NEEDS_OCR` / `SKIPPED_IMAGE` warnings. Format-preserving patches skip edits they cannot apply and report the reason. The HWPX cell-edit path requires the same number of nonempty lines, so blank lines, added or deleted lines, literal `<br>` and ambiguous mappings stay unsupported.

The second tool I lean on at work is [korean-law-mcp](https://github.com/chrisryugj/korean-law-mcp), an MCP server and CLI for Korea's official legal database (the 법제처 Open API). It wraps 42 APIs into 10 tools covering statutes, precedents, administrative rules, local ordinances, treaties and legal interpretations. The part most relevant to LLM work is citation verification: `legal_analysis(mode="verify_citations")` checks the statute articles and case numbers cited in a text against the official source, for existence and for content. It uses kordoc to parse statute annexes, and it is listed in the official MCP Registry. The API key it needs is free.

```
npx --ignore-scripts --omit=optional korean-law-mcp setup
```

If you work with Korean documents, the most useful thing you can send me is a real file that breaks, as an issue. Stars help other people find it too: [github.com/chrisryugj/kordoc](https://github.com/chrisryugj/kordoc).
