{"slug": "nilky-documents-a-floppy-disk-model-failed-tokenizer-and-unfinished-follow-up", "title": "Nilky documents a floppy-disk model, failed tokenizer and unfinished follow-up", "summary": "Hugging Face community contributor Nilky published a September 13th, 2026 retrospective stating that their Single Floppy model, built to fit on one floppy disk, \"scores worse than a 1k model,\" without providing a benchmark method, test set, metric or numerical score. Nilky also reported that a tokenizer failure halted the floppyx3 attempt and that floppyx4 was never finished. The essay supplies no training recipe or results from either follow-up, leaving the comparison as the author's own assessment rather than a reproducible performance result.", "body_md": "# Nilky documents a floppy-disk model, failed tokenizer and unfinished follow-up\n\n**In a September 13th Hugging Face Community Article, the hobbyist says the model performed worse than a 1,000-parameter model, without providing a benchmark method or score.**\n\n        By [RuntimeWire Staff](/author/runtimewire-staff)\n        · Published \n\nPrimary source: [Hugging Face](https://huggingface.co/blog/NILKNARFGonzo/we-got-here)\n\n## Why it matters\n\nTiny-model experiments can expose practical limits that polished release posts omit. Nilky's essay records a poor result, a tokenizer failure and an abandoned follow-up, while its missing benchmark and training details show how little can be concluded from the comparison alone.\n\n[Nilky](https://huggingface.co/NILKNARFGonzo?ref=runtimewire), a [Hugging Face](https://huggingface.co/?ref=runtimewire) community contributor and hobbyist, published a [retrospective account of several tiny language-model experiments](https://huggingface.co/blog/NILKNARFGonzo/we-got-here?ref=runtimewire) on September 13th, 2026. Nilky aimed to make a language model small enough to fit on one floppy disk. In the essay, they say the resulting [Single Floppy model](https://huggingface.co/NILKNARFGonzo/single-floppy-346k?ref=runtimewire) performed worse than a 1,000-parameter model.\n\nThe source is a first-person Community Article. It does not announce a company, funding event or Hugging Face product. Nilky presents the project as a personal experiment shaped by an interest in open-source software and old electronics.\n\nNilky traces that interest in open source to using Raspberry Pi OS. Later experiments involving ChatGPT and DeepSeek led them toward models whose implementations could be inspected directly. Old hardware supplied the storage constraint. \"I really love old electronics,\" Nilky wrote, adding that the essay itself was composed on a 2016 laptop.\n\n### A storage target without a benchmark\n\nNilky calls the result the Single Floppy model and says it \"scores worse than a 1k model.\" The essay supplies no benchmark method, test set, metric or numerical score, so the comparison remains the author's assessment rather than a reproducible performance result.\n\nIt also provides no training recipe or detailed accounting of how the storage constraint affected model quality. Readers cannot determine from the essay which design choice caused the poor result or how the model compares with other tiny language models under equivalent conditions.\n\nThat limited disclosure changes what the experiment can establish. It records the goal and the author's verdict, rather than documenting exactly how a floppy-disk storage limit reduces performance.\n\n### The follow-ups failed differently\n\nThe experiment is useful as a record of constraints rather than a practical model release. [Nilky says](https://huggingface.co/blog/NILKNARFGonzo/we-got-here?ref=runtimewire) a later tokenizer failure stopped the [floppyx3](https://huggingface.co/NILKNARFGonzo/floppyx3-MEGAmodel?ref=runtimewire) attempt, while [floppyx4](https://huggingface.co/NILKNARFGonzo/floppyx4-nonsensicalEssential?ref=runtimewire) was never finished.\n\nThe essay does not explain the tokenizer failure or provide results from either follow-up. Nilky ends by considering the purchase of a dedicated PC for training, another indication that the post is a personal retrospective rather than a maintained development roadmap.\n\nNilky's account is unusually direct about the outcome. The model performed poorly by the creator's own comparison, the next tokenizer failed and a fourth experiment remained incomplete. The useful artifact is the candid record of those constraints and failures, with the technical limits of that record left plainly visible.", "url": "https://wpnews.pro/news/nilky-documents-a-floppy-disk-model-failed-tokenizer-and-unfinished-follow-up", "canonical_source": "https://runtimewire.com/article/nilky-single-floppy-346k-language-model-raspberry-pi", "published_at": "2026-09-13 05:57:33+00:00", "updated_at": "2026-09-13 06:26:50.681858+00:00", "lang": "en", "topics": ["large-language-models", "ai-research"], "entities": ["Nilky", "Hugging Face", "Single Floppy model", "floppyx3", "floppyx4", "Raspberry Pi OS", "ChatGPT", "DeepSeek"], "alternates": {"html": "https://wpnews.pro/news/nilky-documents-a-floppy-disk-model-failed-tokenizer-and-unfinished-follow-up", "markdown": "https://wpnews.pro/news/nilky-documents-a-floppy-disk-model-failed-tokenizer-and-unfinished-follow-up.md", "text": "https://wpnews.pro/news/nilky-documents-a-floppy-disk-model-failed-tokenizer-and-unfinished-follow-up.txt", "jsonld": "https://wpnews.pro/news/nilky-documents-a-floppy-disk-model-failed-tokenizer-and-unfinished-follow-up.jsonld"}}