# VOKU Archive: from paper radio scripts to a 66,010-page Korean dataset (1988–2019)

> Source: <https://discuss.huggingface.co/t/voku-archive-from-paper-radio-scripts-to-a-66-010-page-korean-dataset-1988-2019/180777#post_1>
> Published: 2026-09-28 14:59:21+00:00

Before it was a dataset, VOKU Archive was shelves of paper broadcast cue sheets at a student-run university broadcasting station in South Korea.

Written between 1988 and 2019, these documents preserve 32 years of students preparing campus radio. The project began with a simple aim: to keep that writing readable and accessible beyond the storage room.

Those papers were packed, transported, scanned, and processed. The resulting **66,010-page collection** is now on Hugging Face with page images, OCR text, pseudonym-token annotations, and metadata.

The companion GitHub repository documents the work behind that transformation: AI-assisted processing, human review, corrections, and a runnable demo.

**Dataset:** [NewTypeCola/VOKU-Archive · Datasets at Hugging Face](https://huggingface.co/datasets/NewTypeCola/VOKU-Archive)

**Workflow:** [GitHub - NewTypeCola/VOKU-Archive-Workflow: VOKU Archive Workflow · GitHub](https://github.com/NewTypeCola/VOKU-Archive-Workflow)
