cd /news/large-language-models/coop-a-small-language-model-pretrain… Β· home β€Ί topics β€Ί large-language-models β€Ί article
[ARTICLE Β· art-142798] src=github.com β†— pub= topic=large-language-models verified=true sentiment=↑ positive

Coop: A small language model pretrained by volunteers

A volunteer-run project called Coop has moved to Stage 2, pretraining a ~145M-parameter decoder-only transformer from scratch on FineWeb-Edu using DiLoCo-style low-communication data parallelism, with weights and optimizer state stored only on Hugging Face and aggregation handled by a GitHub Actions cron job. Coop's Stage 1 proof run trained a 15M-parameter model past its Chinchilla-optimal budget on TinyStories in six days, dropping validation loss from 9.01 to 2.8, and the project reports that multiple volunteers on Apple Silicon and plain CPU machines have trained the same outer step and had their updates averaged into one.

read8 min views2 publishedSep 30, 2026
Coop: A small language model pretrained by volunteers
Image: Michielbdejong (auto-discovered)

A small language model pretrained by volunteers. No server, no funding, no daemon β€” the whole training loop runs on donated consumer hardware plus the free tiers of Hugging Face and GitHub Actions.

Stage 2 is live: a ~145M-param model pretraining from scratch on FineWeb-Edu. Stage 1 (15M on TinyStories) completed past its Chinchilla-optimal budget β€” proof that the whole mechanism works. Current numbers and sample output: leaderboard.

The full loop is production-proven, not just designed: multiple volunteers on different machines (Apple Silicon and plain CPU) have trained the same outer step and been averaged into one update β€” the actual data parallelism. A submission that raced a tick was accepted one step later at reduced staleness weight; repeat rounds from one user merged into a single vote (no token farming); coop stop flushed a half-finished round instead of discarding it; and the inbox drains to zero every tick.

DiLoCo-style low-communication data parallelism:

  1. Workers (you) download the current checkpoint from theHF model repo , runH local AdamW steps on a personal data shard, and compute apseudo-gradient :delta = theta_outer - theta_local .
  2. Submission is a pull request against a publicHF dataset repo (the "gradient inbox"), opened withcreate_commit(create_pr=True) . Any free HF account with a write token can submit; the maintainer grants no permissions.
  3. The aggregator is a GitHub Actions cron job (scheduled every 5 min; GitHub's shared scheduler actually fires anywhere from minutes to a few hours apart β€” the protocol tolerates any cadence). Each tick is stateless: it reads the checkpoint and the open inbox PRs, drops over-stale submissions, clips and cosine-gates the rest, robust-aggregates them (trimmed mean / geometric median), takes one Nesterov outer step, uploads the new checkpoint, credits contributors in theledger , regenerates theleaderboard , and closes the processed PRs. Ledger state lives on theledger branch;main only changes through approved pull requests.

Weights and optimizer state live only on Hugging Face (safetensors). Git holds code, config, and the contributor ledger.

Each tick evaluates the new checkpoint on a fixed held-out slice with a fixed seed β€” the same sequences every time, so a change between steps is the model moving and not the eval sampling something else β€” and appends one point to ledger/history.jsonl. A single val loss says nothing: outer steps move it up as often as down. The direction is a property of the series, so the leaderboard and coop status fit a slope over it (against tokens, not outer steps β€” a step is however much work happened to show up that tick) and report it with its standard error, how many steps improved it, and how long it has been since the best one. "Going down" means the slope clears two standard errors; anything less says so instead of pretending.

~145M parameter decoder-only transformer (12 layers, 14 heads, d=896, 1024 context, 32k byte-level BPE vocab, tied embeddings) on FineWeb-Edu β€” real educational web text. GPUs and Apple Silicon pull their weight here; plain CPUs are better suited to CPU-tier work (see CONTRIBUTING.md).

The proof run: a 15M-param model pretrained past its Chinchilla-optimal budget on TinyStories by volunteers in six days, val loss 9.01 β†’ 2.8. It stays usable forever β€” commonsense-ai/tinystories-15m has the weights, the model card, and a working load-and-generate snippet. The final stage-1 leaderboard is archived as LEADERBOARD-stage1.md on the ledger branch.

Talk to whatever the volunteers have trained so far β€” no account, no clone, no training:

npx coop-ai run latest     # bun: bunx coop-ai run latest

It downloads the current checkpoint (cached after the first time), then takes prompts and writes what comes next. run tinystories plays the finished stage-1 model instead, which is far more coherent than a run still in progress. One-off and pipeable:

coop run latest --prompt "The best way to learn mathematics is"
echo "Once upon a time" | coop run tinystories

--tokens, --temperature, --top-k, and --device are there when you want them; --revision pins a specific checkpoint.

One command if you have Node or Bun (it uses uv under the hood and tells you how to get it):

npx coop-ai start     # bun: bunx coop-ai start

Or install with uv directly:

uv tool install git+https://github.com/commonsense-ai/coop
coop start

(uv tool install coop-ai / pipx install coop-ai once the PyPI package clears review.)

The first run asks you to paste a Hugging Face write token (free account) β€” after that it's zero-setup. The worker runs in the background: it builds you a personal data shard (a slice derived from your username so volunteers don't overlap), then trains and submits rounds until you say otherwise.

coop start then opens the live progress screen, so you can watch the one-time prep and the first round go by:

coop Β· fineweb-150m Β· naloxene Β· Apple GPU

this round   β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘  51.2%  inner step 256/500 Β· loss 3.21 Β· ~2m 02s left
the model    β–ˆβ–ˆβ–ˆβ–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘  12.4%  37.2M of ~300.0M tokens
your share   β–ˆβ–ˆβ–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘   6.6%  rank 2 of 14 Β· 2,457,600 tokens

3 rounds Β· 2,457,600 tokens this session β€” each one submitted for you

  > keep training (leave this screen)
    stop contributing

↑↓ move Β· enter choose Β· ←→ advanced view Β· q leave (training keeps going)

Left/right swaps the simple view for the advanced one (every field coop status prints). Up/down and enter stop the worker without leaving the screen. q leaves the screen and keeps training β€” closing the view never stops anything.

coop progress            # reopen it any time (`--advanced` starts on the detail view)
coop progress --once     # one snapshot, no screen; good for pipes
coop progress --auto off # `coop start` goes back to printing a summary
coop status    # the same facts as one printout
coop logs -f   # watch it work
coop stop      # stop contributing; `coop start` resumes any time

coop start --rounds 3 contributes a fixed number of rounds and stops by itself, and coop start --no-progress skips the screen just this once.

Every release publishes release.json to the ledger branch, and that is the only thing a volunteer's machine polls. coop status says when a newer version is out, and:

coop update            # get it now (works out how you installed coop)
coop update --check    # what's new, install nothing
coop update --auto on  # keep it current by itself

With --auto on, a running worker adopts the new version between rounds β€” never mid-round, so a restart can't cost you trained work β€” and picks its training back up where it left off. Off by default: nothing on your machine changes unless you ask. A clone is left alone either way; there, git pull is the update.

One exception, and it only fires when your worker is already broken. If it cannot finish a single round β€” several failures in a row, a restart, still nothing β€” it will take a newer version if one exists, even with auto-update off. You asked it to contribute, it isn't contributing, and a fix on the channel is the only thing left that can change that. coop status says so, and coop logs shows the attempt.

Self-updating only works from 0.3.0 on, so an install older than that can't reach it:

coop: error: argument cmd: invalid choice: 'update'

That means the coop on your PATH predates the command. Reinstall it once β€” uv tool install --force git+https://github.com/commonsense-ai/coop, or just use npx coop-ai, which always resolves the current code β€” and it keeps itself current from then on. uv tool list is worth a look if you installed early: the package was once named coop rather than coop-ai, and a leftover of the old name claims the same coop executable.

Prefer a foreground one-off? uvx --from git+https://github.com/commonsense-ai/coop coop-join --hf-token hf_... runs rounds until ctrl-c (--once for a single round, --device cuda|mps|tpu|cpu to override).

coop trains on one when torch_xla is importable and it reports a real TPU β€” detection is automatic, and rounds land on the leaderboard at tpu tier. The catch is the environment: torch_xla is version-locked to torch, so it has to be installed into the same env as coop. From a clone (uv pip install torch_xla matched to your torch, then uv run python -m coop.trainer --data ... --loop) you control both, which is why that's the route we'd suggest on a TPU VM. uvx --with torch_xla --from git+https://github.com/commonsense-ai/coop coop-join is the one-liner shape if you'd rather; whether it resolves depends on your torch.

Two things to know before you spend a TPU on this. A single worker process uses one chip β€” on a v5e-8 that's an eighth of the board, since coop doesn't spawn per-core replicas. And at 15M parameters the steps are small enough that XLA compilation and host overhead eat much of the advantage, so a mid-range GPU is likely to out-earn a TPU here. Playing with the model (coop run) deliberately stays on the CPU: XLA recompiles for every new sequence length, which makes generation slower on a TPU than off it.

From a clone, the equivalent is:

uv sync   # NVIDIA box? add `uv pip install --torch-backend auto torch`:
export HF_TOKEN=hf_...
uv run python -m coop.data --skip 0 --docs 20000
uv run python -m coop.trainer --data data/shard_0_20000.bin --loop

Run it as often as you like. Accepted submissions earn tokens on the leaderboard; CPU-only machines can also contribute tokenization, dedup, filtering, and eval runs (see CONTRIBUTING.md).

export HF_TOKEN=hf_...   # write access to both HF repos
uv run python -m coop.data --train-tokenizer --skip 0 --docs 50000
uv run python -m coop.aggregate --init          # genesis checkpoint at step 0

Then set the HF_TOKEN repo secret on GitHub so .github/workflows/aggregate.yml can run the tick.

uv run pytest -q
uv run ruff check .

Everything is configured in config/run.yaml. Architecture rules live in AGENTS.md.

License: undecided β€” all rights reserved for now, see LICENSE.md.

── more in #large-language-models 4 stories Β· sorted by recency
── more on @coop 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/coop-a-small-languag…] indexed:0 read:8min 2026-09-30 Β· β€”