cd /news/large-language-models/nilky-documents-a-floppy-disk-model-… · home topics large-language-models article
[ARTICLE · art-128106] src=runtimewire.com ↗ pub= topic=large-language-models verified=true sentiment=↓ negative

Nilky documents a floppy-disk model, failed tokenizer and unfinished follow-up

Hugging Face community contributor Nilky published a September 13th, 2026 retrospective stating that their Single Floppy model, built to fit on one floppy disk, "scores worse than a 1k model," without providing a benchmark method, test set, metric or numerical score. Nilky also reported that a tokenizer failure halted the floppyx3 attempt and that floppyx4 was never finished. The essay supplies no training recipe or results from either follow-up, leaving the comparison as the author's own assessment rather than a reproducible performance result.

read2 min views2 publishedSep 13, 2026
Nilky documents a floppy-disk model, failed tokenizer and unfinished follow-up
Image: Runtimewire (auto-discovered)

In a September 13th Hugging Face Community Article, the hobbyist says the model performed worse than a 1,000-parameter model, without providing a benchmark method or score.

        By [RuntimeWire Staff](/author/runtimewire-staff)
        · Published 

Primary source: [Hugging Face](https://huggingface.co/blog/NILKNARFGonzo/we-got-here)

Why it matters #

Tiny-model experiments can expose practical limits that polished release posts omit. Nilky's essay records a poor result, a tokenizer failure and an abandoned follow-up, while its missing benchmark and training details show how little can be concluded from the comparison alone.

Nilky, a Hugging Face community contributor and hobbyist, published a retrospective account of several tiny language-model experiments on September 13th, 2026. Nilky aimed to make a language model small enough to fit on one floppy disk. In the essay, they say the resulting Single Floppy model performed worse than a 1,000-parameter model.

The source is a first-person Community Article. It does not announce a company, funding event or Hugging Face product. Nilky presents the project as a personal experiment shaped by an interest in open-source software and old electronics.

Nilky traces that interest in open source to using Raspberry Pi OS. Later experiments involving ChatGPT and DeepSeek led them toward models whose implementations could be inspected directly. Old hardware supplied the storage constraint. "I really love old electronics," Nilky wrote, adding that the essay itself was composed on a 2016 laptop.

A storage target without a benchmark

Nilky calls the result the Single Floppy model and says it "scores worse than a 1k model." The essay supplies no benchmark method, test set, metric or numerical score, so the comparison remains the author's assessment rather than a reproducible performance result.

It also provides no training recipe or detailed accounting of how the storage constraint affected model quality. Readers cannot determine from the essay which design choice caused the poor result or how the model compares with other tiny language models under equivalent conditions.

That limited disclosure changes what the experiment can establish. It records the goal and the author's verdict, rather than documenting exactly how a floppy-disk storage limit reduces performance.

The follow-ups failed differently

The experiment is useful as a record of constraints rather than a practical model release. Nilky says a later tokenizer failure stopped the floppyx3 attempt, while floppyx4 was never finished.

The essay does not explain the tokenizer failure or provide results from either follow-up. Nilky ends by considering the purchase of a dedicated PC for training, another indication that the post is a personal retrospective rather than a maintained development roadmap.

Nilky's account is unusually direct about the outcome. The model performed poorly by the creator's own comparison, the next tokenizer failed and a fourth experiment remained incomplete. The useful artifact is the candid record of those constraints and failures, with the technical limits of that record left plainly visible.

── more in #large-language-models 4 stories · sorted by recency
── more on @nilky 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/nilky-documents-a-fl…] indexed:0 read:2min 2026-09-13 ·