Time Capsule of Testable Human Knowledge: 41 Years of Jeopardy! in a Single Free Local Model A 9 GB 4-bit quantized Qwen2.5-14B model running locally answers 67% of the full 529,939-clue Jeopardy! corpus and 65% on post-cutoff clues, according to a new arXiv paper (2608.27459). The benchmark shows mid-size local models can handle broad factual recall offline for free, though they trail frontier models like Claude Opus, which scores 95%. arXiv https://arxiv.org/abs/2608.27459 Time Capsule of Testable Human Knowledge: 41 Years of Jeopardy in a Single Free Local Model Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated. A 9 GB 4-bit Qwen2.5-14B running locally answers 67% of the full 529,939-clue Jeopardy corpus and 65% on post-cutoff clues, meaning a broad factual-recall capability that once required a server cluster now fits in a file you can run offline for free. For anyone shipping, this quantifies that mid-size local models are viable for general-knowledge lookup without API dependency—but the 65% vs. 95% gap against frontier models Claude Opus is your accuracy tax for going local, so reserve small local models for latency/cost-sensitive recall and keep frontier calls for high-stakes factoid precision. Mistral Large quota or rate limit — check usage and plan. Original headline: Time Capsule of Testable Human Knowledge: 41 Years of Jeopardy in a Single Free Local Model