{"slug": "coop-a-small-language-model-pretrained-by-volunteers", "title": "Coop: A small language model pretrained by volunteers", "summary": "A volunteer-run project called Coop has moved to Stage 2, pretraining a ~145M-parameter decoder-only transformer from scratch on FineWeb-Edu using DiLoCo-style low-communication data parallelism, with weights and optimizer state stored only on Hugging Face and aggregation handled by a GitHub Actions cron job. Coop's Stage 1 proof run trained a 15M-parameter model past its Chinchilla-optimal budget on TinyStories in six days, dropping validation loss from 9.01 to 2.8, and the project reports that multiple volunteers on Apple Silicon and plain CPU machines have trained the same outer step and had their updates averaged into one.", "body_md": "A small language model pretrained by volunteers. No server, no funding, no daemon — the whole training loop runs on donated consumer hardware plus the free tiers of Hugging Face and GitHub Actions.\n\n**Stage 2 is live**: a ~145M-param model pretraining from scratch on FineWeb-Edu.\nStage 1 (15M on TinyStories) completed past its Chinchilla-optimal budget — proof\nthat the whole mechanism works. Current numbers and sample output:\n[leaderboard](https://github.com/commonsense-ai/coop/blob/ledger/LEADERBOARD.md).\n\nThe full loop is production-proven, not just designed: multiple volunteers on\ndifferent machines (Apple Silicon and plain CPU) have trained the same outer step\nand been averaged into one update — the actual data parallelism. A submission that\nraced a tick was accepted one step later at reduced staleness weight; repeat rounds\nfrom one user merged into a single vote (no token farming); `coop stop` flushed a\nhalf-finished round instead of discarding it; and the inbox drains to zero every\ntick.\n\nDiLoCo-style low-communication data parallelism:\n\n1. **Workers** (you) download the current checkpoint from the[HF model repo](https://huggingface.co/commonsense-ai/fineweb-150m) ,\nrun`H` local AdamW steps on a personal data shard, and compute a**pseudo-gradient** :`delta = theta_outer - theta_local` .\n2. **Submission** is a pull request against a public[HF dataset repo](https://huggingface.co/datasets/commonsense-ai/fineweb-150m-inbox) (the \"gradient inbox\"), opened with`create_commit(create_pr=True)` . Any free HF\naccount with a write token can submit; the maintainer grants no permissions.\n3. **The aggregator** is a GitHub Actions cron job (scheduled every 5 min;\nGitHub's shared scheduler actually fires anywhere from minutes to a few hours\napart — the protocol tolerates any cadence). Each tick is stateless: it reads the checkpoint and the open inbox PRs, drops over-stale\nsubmissions, clips and cosine-gates the rest, robust-aggregates them\n(trimmed mean / geometric median), takes one Nesterov outer step, uploads the\nnew checkpoint, credits contributors in the[ledger](https://github.com/commonsense-ai/coop/tree/ledger/ledger) ,\nregenerates the[leaderboard](https://github.com/commonsense-ai/coop/blob/ledger/LEADERBOARD.md) ,\nand closes the processed PRs. Ledger state lives on the`ledger` branch;`main` only changes through approved pull requests.\n\nWeights and optimizer state live **only** on Hugging Face (safetensors). Git holds\ncode, config, and the contributor ledger.\n\nEach tick evaluates the new checkpoint on a fixed held-out slice with a fixed seed —\nthe same sequences every time, so a change between steps is the model moving and not\nthe eval sampling something else — and appends one point to\n[`ledger/history.jsonl`](https://github.com/commonsense-ai/coop/blob/ledger/ledger/history.jsonl).\nA single val loss says nothing: outer steps move it up as often as down. The direction\nis a property of the series, so the leaderboard and `coop status` fit a slope over it\n(against tokens, not outer steps — a step is however much work happened to show up that\ntick) and report it with its standard error, how many steps improved it, and how long\nit has been since the best one. \"Going down\" means the slope clears two standard errors;\nanything less says so instead of pretending.\n\n~145M parameter decoder-only transformer (12 layers, 14 heads, d=896, 1024 context,\n32k byte-level BPE vocab, tied embeddings) on\n[FineWeb-Edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu) — real\neducational web text. GPUs and Apple Silicon pull their weight here; plain CPUs are\nbetter suited to CPU-tier work (see [CONTRIBUTING.md](https://github.com/commonsense-ai/coop/blob/main/CONTRIBUTING.md)).\n\nThe proof run: a 15M-param model pretrained past its Chinchilla-optimal budget on\nTinyStories by volunteers in six days, val loss 9.01 → 2.8. It stays usable forever —\n[commonsense-ai/tinystories-15m](https://huggingface.co/commonsense-ai/tinystories-15m)\nhas the weights, the model card, and a working load-and-generate snippet. The final\nstage-1 leaderboard is archived as `LEADERBOARD-stage1.md` on the ledger branch.\n\nTalk to whatever the volunteers have trained so far — no account, no clone, no training:\n\n```\nnpx coop-ai run latest     # bun: bunx coop-ai run latest\n```\n\nIt downloads the current checkpoint (cached after the first time), then takes\nprompts and writes what comes next. `run tinystories` plays the finished stage-1\nmodel instead, which is far more coherent than a run still in progress. One-off\nand pipeable:\n\n```\ncoop run latest --prompt \"The best way to learn mathematics is\"\necho \"Once upon a time\" | coop run tinystories\n```\n\n`--tokens`, `--temperature`, `--top-k`, and `--device` are there when you want\nthem; `--revision` pins a specific checkpoint.\n\nOne command if you have Node or [Bun](https://bun.sh) (it uses\n[uv](https://docs.astral.sh/uv/) under the hood and tells you how to get it):\n\n```\nnpx coop-ai start     # bun: bunx coop-ai start\n```\n\nOr install with uv directly:\n\n```\nuv tool install git+https://github.com/commonsense-ai/coop\ncoop start\n```\n\n(`uv tool install coop-ai` / `pipx install coop-ai` once the PyPI package\nclears review.)\n\nThe first run asks you to paste a Hugging Face\n[write token](https://huggingface.co/settings/tokens) (free account) — after that\nit's zero-setup. The worker runs in the background: it builds you a personal data\nshard (a slice derived from your username so volunteers don't overlap), then trains\nand submits rounds until you say otherwise.\n\n`coop start` then opens the live progress screen, so you can watch the one-time prep\nand the first round go by:\n\n```\ncoop · fineweb-150m · naloxene · Apple GPU\n\nthis round   ████████████░░░░░░░░░░░░  51.2%  inner step 256/500 · loss 3.21 · ~2m 02s left\nthe model    ███░░░░░░░░░░░░░░░░░░░░░  12.4%  37.2M of ~300.0M tokens\nyour share   ██░░░░░░░░░░░░░░░░░░░░░░   6.6%  rank 2 of 14 · 2,457,600 tokens\n\n3 rounds · 2,457,600 tokens this session — each one submitted for you\n\n  > keep training (leave this screen)\n    stop contributing\n\n↑↓ move · enter choose · ←→ advanced view · q leave (training keeps going)\n```\n\nLeft/right swaps the simple view for the advanced one (every field `coop status`\nprints). Up/down and enter stop the worker without leaving the screen. `q` leaves the\nscreen and keeps training — closing the view never stops anything.\n\n```\ncoop progress            # reopen it any time (`--advanced` starts on the detail view)\ncoop progress --once     # one snapshot, no screen; good for pipes\ncoop progress --auto off # `coop start` goes back to printing a summary\ncoop status    # the same facts as one printout\ncoop logs -f   # watch it work\ncoop stop      # stop contributing; `coop start` resumes any time\n```\n\n`coop start --rounds 3` contributes a fixed number of rounds and stops by itself, and\n`coop start --no-progress` skips the screen just this once.\n\nEvery release publishes `release.json` to the `ledger` branch, and that is the only\nthing a volunteer's machine polls. `coop status` says when a newer version is out,\nand:\n\n```\ncoop update            # get it now (works out how you installed coop)\ncoop update --check    # what's new, install nothing\ncoop update --auto on  # keep it current by itself\n```\n\nWith `--auto on`, a running worker adopts the new version **between rounds** — never\nmid-round, so a restart can't cost you trained work — and picks its training back up\nwhere it left off. Off by default: nothing on your machine changes unless you ask.\nA clone is left alone either way; there, `git pull` is the update.\n\nOne exception, and it only fires when your worker is already broken. If it cannot\nfinish a single round — several failures in a row, a restart, still nothing — it will\ntake a newer version if one exists, even with auto-update off. You asked it to\ncontribute, it isn't contributing, and a fix on the channel is the only thing left\nthat can change that. `coop status` says so, and `coop logs` shows the attempt.\n\nSelf-updating only works from 0.3.0 on, so an install older than that can't reach it:\n\n```\ncoop: error: argument cmd: invalid choice: 'update'\n```\n\nThat means the `coop` on your PATH predates the command. Reinstall it once —\n`uv tool install --force git+https://github.com/commonsense-ai/coop`, or just use\n`npx coop-ai`, which always resolves the current code — and it keeps itself current\nfrom then on. `uv tool list` is worth a look if you installed early: the package was\nonce named `coop` rather than `coop-ai`, and a leftover of the old name claims the\nsame `coop` executable.\n\nPrefer a foreground one-off? `uvx --from git+https://github.com/commonsense-ai/coop coop-join --hf-token hf_...` runs rounds until ctrl-c (`--once` for a single round,\n`--device cuda|mps|tpu|cpu` to override).\n\ncoop trains on one when `torch_xla` is importable and it reports a real TPU —\ndetection is automatic, and rounds land on the leaderboard at `tpu` tier. The\ncatch is the environment: `torch_xla` is version-locked to torch, so it has to be\ninstalled into the *same* env as coop. From a clone (`uv pip install torch_xla`\nmatched to your torch, then `uv run python -m coop.trainer --data ... --loop`) you\ncontrol both, which is why that's the route we'd suggest on a TPU VM. `uvx --with torch_xla --from git+https://github.com/commonsense-ai/coop coop-join` is the\none-liner shape if you'd rather; whether it resolves depends on your torch.\n\nTwo things to know before you spend a TPU on this. A single worker process uses one\nchip — on a v5e-8 that's an eighth of the board, since coop doesn't spawn per-core\nreplicas. And at 15M parameters the steps are small enough that XLA compilation and\nhost overhead eat much of the advantage, so a mid-range GPU is likely to out-earn a\nTPU here. Playing with the model (`coop run`) deliberately stays on the CPU: XLA\nrecompiles for every new sequence length, which makes generation slower on a TPU\nthan off it.\n\nFrom a clone, the equivalent is:\n\n```\nuv sync   # NVIDIA box? add `uv pip install --torch-backend auto torch`:\n          # the lockfile ships CPU wheels so aggregator ticks stay fast\nexport HF_TOKEN=hf_...\nuv run python -m coop.data --skip 0 --docs 20000\nuv run python -m coop.trainer --data data/shard_0_20000.bin --loop\n```\n\nRun it as often as you like. Accepted submissions earn tokens on the leaderboard;\nCPU-only machines can also contribute tokenization, dedup, filtering, and eval runs\n(see [CONTRIBUTING.md](https://github.com/commonsense-ai/coop/blob/main/CONTRIBUTING.md)).\n\n```\nexport HF_TOKEN=hf_...   # write access to both HF repos\nuv run python -m coop.data --train-tokenizer --skip 0 --docs 50000\nuv run python -m coop.aggregate --init          # genesis checkpoint at step 0\n```\n\nThen set the `HF_TOKEN` repo secret on GitHub so `.github/workflows/aggregate.yml`\ncan run the tick.\n\n```\nuv run pytest -q\nuv run ruff check .\n```\n\nEverything is configured in [config/run.yaml](https://github.com/commonsense-ai/coop/blob/main/config/run.yaml). Architecture rules\nlive in [AGENTS.md](https://github.com/commonsense-ai/coop/blob/main/AGENTS.md).\n\nLicense: undecided — all rights reserved for now, see [LICENSE.md](https://github.com/commonsense-ai/coop/blob/main/LICENSE.md).", "url": "https://wpnews.pro/news/coop-a-small-language-model-pretrained-by-volunteers", "canonical_source": "https://github.com/commonsense-ai/coop", "published_at": "2026-09-30 19:51:50+00:00", "updated_at": "2026-09-30 20:19:56.488290+00:00", "lang": "en", "topics": ["large-language-models", "machine-learning", "ai-research", "ai-infrastructure", "mlops"], "entities": ["Coop", "Hugging Face", "GitHub Actions", "FineWeb-Edu", "TinyStories", "commonsense-ai/tinystories-15m", "commonsense-ai/fineweb-150m", "DiLoCo"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/coop-a-small-language-model-pretrained-by-volunteers", "markdown": "https://wpnews.pro/news/coop-a-small-language-model-pretrained-by-volunteers.md", "text": "https://wpnews.pro/news/coop-a-small-language-model-pretrained-by-volunteers.txt", "jsonld": "https://wpnews.pro/news/coop-a-small-language-model-pretrained-by-volunteers.jsonld"}}