cd /news/artificial-intelligence/complete-personal-archive-of-tsiolko… · home topics artificial-intelligence article
[ARTICLE · art-109672] src=discuss.huggingface.co ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Complete personal archive of Tsiolkovsky (51k sheets) — with recognition accuracy actually measured

Vladimir Besk released the complete personal archive of Konstantin Tsiolkovsky as a dataset on Hugging Face, comprising 51,008 sheets across 2,019 archival files under CC0. The archive includes 1,759 manuscript-typescript pairs from 224 files, where two readings of the same text agree on a median 37% of words, and against printed editions, recognition accuracy is 98.1% on typescript and 81.1% on handwriting. This dataset provides a rare opportunity for handwriting recognition research, including fine-tuning models for Russian handwriting and testing uncertainty calibration.

read3 min views1 publishedAug 25, 2026

I’ve released the complete personal archive of Konstantin Tsiolkovsky as a dataset: 51,008 sheets, 2,019 archival files, CC0.

vladimirbesk/tsiolkovsky-papers · Datasets at Hugging Face Search it in your browser: Tsiolkovsky Papers — full-text search - a Hugging Face Space by vladimirbesk

Who he was. A provincial schoolteacher in Kaluga, deaf since a childhood illness, largely self-taught. In 1897 he derived what we now call the rocket equation — the relation between exhaust velocity, mass ratio and the speed a rocket can reach. He described multistage rockets, airlocks, closed-loop life support and the space elevator decades before anyone could test any of it. He also spent thirty years on all-metal airships, which is most of what the archive actually contains: the man who is remembered for rockets wrote more about dirigibles.

He died in 1935. His papers went to the Archive of the Russian Academy of Sciences as fond 555, were scanned, and were published online — page by page, with no catalogue you could query, no full-text search, and no dataset. Open, and unusable.

Why it might interest you as data.

One hand, fifty years. Nearly everything here is written by the same person between the 1880s and 1935, which is unusual: most handwriting corpora are many writers, shallow per writer. Here you can watch one hand age.

An orthographic break in the middle. Russian spelling was reformed in 1918, and the archive spans it — pre-reform ѣ, і, ъ, ѳ in the early material, modern spelling later, and inconsistency in between. Hard for a model, interesting for anyone working on historical language.

Not just prose. Derivations, tables, engineering drawings with captions, marginalia, and a great deal of authorial revision: 36,446 struck-out passages are marked as such.

The part I’d point at. Transcription accuracy here is measured, which for handwritten archives it usually is not — measuring normally requires ground truth the archive doesn’t have. This one does, accidentally: part of the fond holds each text twice, as the author’s manuscript and as a typed copy of it, filed together. Read both, and the disagreement is almost purely the difficulty of the hand, because the text and the pipeline are identical.

Over 1,759 such pairs from 224 files, two readings of one and the same text agree on a median 37% of words. Against printed editions: 98.1% of characters on typescript, 81.1% on handwriting.

That 37% is the number I’d hand to anyone who wants to know what state handwritten Russian recognition is actually in.

Things you could do with it. There is, as far as I can find, no HTR model for Russian handwriting on the Hub at all — nothing to fine-tune from, and now some data to fine-tune on. The corpus also carries 299,939 explicit uncertainty marks the model left on itself, which makes it a ready testbed for uncertainty calibration: the marks are the model’s own doubt, and the manuscript/typescript pairs are an external check on whether that doubt was warranted. And 376 pairs of files share verbatim text — the same work rewritten under different archival titles — if text reuse is your thing.

What it will not bear. No human has checked this against the scans. Not one sheet. It is a machine transcription with its doubts marked, good for search, retrieval and training, not for quotation without checking the scan — every row carries a link to the original image. And redactions cannot be collated at this quality: two variants of one work share 19% of their words, below the rate at which two readings of a single page agree, so authorial revision cannot be told from misreading.

Method: arXiv:2608.03617. Archive of record: 10.5281/zenodo.21705221. Everything CC0 — take it without asking.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @vladimir besk 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/complete-personal-ar…] indexed:0 read:3min 2026-08-25 ·