VOKU Archive: from paper radio scripts to a 66,010-page Korean dataset (1988–2019) The VOKU Archive, a 66,010-page dataset of Korean campus radio scripts written between 1988 and 2019, has been published on Hugging Face by the creator NewTypeCola with page images, OCR text, pseudonym-token annotations, and metadata. A companion GitHub repository, VOKU-Archive-Workflow, documents the AI-assisted processing, human review, corrections, and a runnable demo behind the digitization of the paper broadcast cue sheets from a student-run South Korean university station. Before it was a dataset, VOKU Archive was shelves of paper broadcast cue sheets at a student-run university broadcasting station in South Korea. Written between 1988 and 2019, these documents preserve 32 years of students preparing campus radio. The project began with a simple aim: to keep that writing readable and accessible beyond the storage room. Those papers were packed, transported, scanned, and processed. The resulting 66,010-page collection is now on Hugging Face with page images, OCR text, pseudonym-token annotations, and metadata. The companion GitHub repository documents the work behind that transformation: AI-assisted processing, human review, corrections, and a runnable demo. Dataset: NewTypeCola/VOKU-Archive · Datasets at Hugging Face https://huggingface.co/datasets/NewTypeCola/VOKU-Archive Workflow: GitHub - NewTypeCola/VOKU-Archive-Workflow: VOKU Archive Workflow · GitHub https://github.com/NewTypeCola/VOKU-Archive-Workflow