TL;DR: VIDRAFT's Darwin-180B-RSI is a 180-billion-parameter reasoning model trained via a Recursive Self-Improvement (RSI) framework that updates model weights directly — no human-annotated chain-of-thought required. It just claimed #1 on both the LEXam and LEXam-hard Hugging Face-certified legal reasoning benchmarks, despite never being trained on legal domain data. Its checkpoint is publicly available on Hugging Face for external reproduction.
Darwin-180B-RSI is VIDRAFT's self-improving large language model, built around a model-level Recursive Self-Improvement (RSI) architecture. Key facts from the announcement:
The core design principle is Recursive Self-Improvement at the weight level — a conceptually distinct approach from system-level self-improvement (like prompt tuning or tool-orchestration loops that leave base weights unchanged).
Here's the high-level loop:
The practical result is that reasoning capabilities generalized from math and science transfer to unseen domains like legal analysis — the model isn't memorizing legal statutes, it's applying structured multi-step logical inference. As a secondary efficiency gain, the RSI-trained model reportedly shortened reasoning token length by 11% relative to the base model, pruning redundant inference steps without sacrificing accuracy.
This is architecturally different from Google's system-level RRSI approach (which adjusts prompts, external tools, and workflows while keeping base weights frozen). VIDRAFT's approach bakes improvements permanently into a single weight file, meaning users get the full capability from the model artifact alone, with no additional scaffolding required at inference time.
All scores below are from the Hugging Face-certified LEXam benchmark, developed jointly by ETH Zurich, University of Zurich, and the Max Planck Institute. LEXam is built from 340 real law school final exams and tests applied legal reasoning under civil law — not statute memorization.
LEXam (multiple-choice, 1,655 questions):
| Model | Score |
|---|---|
| Darwin-180B-RSI (4-pass majority vote) | 68.94 |
| Darwin-180B-RSI (1-pass, single inference) | 60.54 |
| GPT-5 (benchmark authors' measurement) | 62.65 |
| Claude-4.5-Sonnet (benchmark authors' measurement) | 58.01 |
| Gemini-2.5-Pro (benchmark authors' measurement) | 55.72 |
| DeepSeek-R1 (685B) | 52.41 |
| Qwen3-235B-Thinking | 48.19 |
| gpt-oss-120b | 47.71 |
LEXam-hard (open-ended, 518 questions — scored by official grading model):
| Model | Score |
|---|---|
| Darwin-180B-RSI | 45.72 |
| Inkling (Thinking Machines) | 40.82 |
| DeepSeek-V4-Pro | 38.93 |
Context: random-chance baseline on the 4-option multiple-choice LEXam is 25 points. Most top LLMs average ~50 points on it; LEXam-hard's previous best was in the low 40s.
Across all 47 Hugging Face certified leaderboards, VIDRAFT now holds #1 in 7 categories — ahead of Zhipu AI (4), DeepSeek / Xiaomi / Moonshot AI (3 each).
The Darwin-180B-RSI model checkpoint and its evaluation environment are publicly available on Hugging Face as open-source. VIDRAFT has also disclosed the full benchmark evaluation parameters — number of inference passes, ensemble majority-vote settings, and reasoning token budget — to enable independent reproduction.
To pull the model via the Hugging Face CLI:
huggingface-cli download VIDRAFT/Darwin-180B-RSI
⚠️ Verify the exact repository path on huggingface.co/VIDRAFT before down. The source confirms open availability but does not specify a direct URL string.
Q: Does Darwin-180B-RSI require legal fine-tuning data to perform well on legal benchmarks?
A: No. The model was explicitly evaluated without any prior legal domain training. Its legal reasoning performance comes from generalizing multi-step logical inference skills acquired during math and science self-improvement cycles.
Q: What's the practical difference between VIDRAFT's model-level RSI and system-level self-improvement (like Google's RRSI)?
A: System-level approaches adjust prompts, tools, and orchestration workflows at runtime while leaving the underlying model weights unchanged. VIDRAFT's RSI updates the model's weights directly, so improvements are permanently embedded in the weight file. You don't need any additional runtime infrastructure — just load the model and run inference. VIDRAFT notes both approaches are complementary rather than competing.
Q: How do I know the benchmark results are legitimate and not cherry-picked evaluation conditions?
A: VIDRAFT publicly disclosed all evaluation parameters: inference pass count, majority-vote configuration, and reasoning token budget. The checkpoint is on Hugging Face so anyone can re-run the LEXam evaluation independently. The benchmark itself was built by academic researchers at ETH Zurich, University of Zurich, and the Max Planck Institute, with answers validated by legal professionals.
Originally reported by 로이슈 (2026-10-01) — source article.