{"slug": "voku-archive-from-paper-radio-scripts-to-a-66010-page-korean-dataset-1988-2019", "title": "VOKU Archive: from paper radio scripts to a 66,010-page Korean dataset (1988–2019)", "summary": "The VOKU Archive, a 66,010-page dataset of Korean campus radio scripts written between 1988 and 2019, has been published on Hugging Face by the creator NewTypeCola with page images, OCR text, pseudonym-token annotations, and metadata. A companion GitHub repository, VOKU-Archive-Workflow, documents the AI-assisted processing, human review, corrections, and a runnable demo behind the digitization of the paper broadcast cue sheets from a student-run South Korean university station.", "body_md": "Before it was a dataset, VOKU Archive was shelves of paper broadcast cue sheets at a student-run university broadcasting station in South Korea.\n\nWritten between 1988 and 2019, these documents preserve 32 years of students preparing campus radio. The project began with a simple aim: to keep that writing readable and accessible beyond the storage room.\n\nThose papers were packed, transported, scanned, and processed. The resulting **66,010-page collection** is now on Hugging Face with page images, OCR text, pseudonym-token annotations, and metadata.\n\nThe companion GitHub repository documents the work behind that transformation: AI-assisted processing, human review, corrections, and a runnable demo.\n\n**Dataset:** [NewTypeCola/VOKU-Archive · Datasets at Hugging Face](https://huggingface.co/datasets/NewTypeCola/VOKU-Archive)\n\n**Workflow:** [GitHub - NewTypeCola/VOKU-Archive-Workflow: VOKU Archive Workflow · GitHub](https://github.com/NewTypeCola/VOKU-Archive-Workflow)", "url": "https://wpnews.pro/news/voku-archive-from-paper-radio-scripts-to-a-66010-page-korean-dataset-1988-2019", "canonical_source": "https://discuss.huggingface.co/t/voku-archive-from-paper-radio-scripts-to-a-66-010-page-korean-dataset-1988-2019/180777#post_1", "published_at": "2026-09-28 14:59:21+00:00", "updated_at": "2026-09-28 15:21:51.639183+00:00", "lang": "en", "topics": ["artificial-intelligence", "natural-language-processing", "ai-research"], "entities": ["VOKU Archive", "Hugging Face", "GitHub", "NewTypeCola", "VOKU-Archive-Workflow"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/voku-archive-from-paper-radio-scripts-to-a-66010-page-korean-dataset-1988-2019", "markdown": "https://wpnews.pro/news/voku-archive-from-paper-radio-scripts-to-a-66010-page-korean-dataset-1988-2019.md", "text": "https://wpnews.pro/news/voku-archive-from-paper-radio-scripts-to-a-66010-page-korean-dataset-1988-2019.txt", "jsonld": "https://wpnews.pro/news/voku-archive-from-paper-radio-scripts-to-a-66010-page-korean-dataset-1988-2019.jsonld"}}