cd /news/artificial-intelligence/shelf-a-synthetic-harness-for-multi-… · home topics artificial-intelligence article
[ARTICLE · art-121086] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

SHELF: A Synthetic Harness for Multi-Task Bibliographic Benchmarking

Researchers released SHELF, a Python benchmark system for evaluating large language models on library and archive tasks, generating 62,899 model-written documents from Library of Congress vocabularies. Subject classification reached 0.8887, while genre-form classification lagged at 0.2605, with sparse methods like TF-IDF remaining competitive and fastest. The system, available on GitHub and Hugging Face, enables controlled benchmarking and generation of unseen documents, though scores do not predict production catalogue accuracy.

read1 min views2 publishedSep 4, 2026

arXiv:2609.03047v1 Announce Type: new Abstract: Libraries and archives manage large collections with limited staff and computing budgets, yet common benchmarks do not systematically test their bibliographic work. They need to know which methods work for their tasks and what those methods require to run. SHELF, the Synthetic Harness for Evaluating LLM Fitness, addresses this gap. It is a Python system that turns labelled taxonomies, writing specifications, and a generation budget into controlled benchmark data and evaluation tasks. This first release contains 62,899 model-written documents based on Library of Congress vocabularies, with tasks for classification, clustering, retrieval, pair classification, and instruction retrieval. We compare TF, TF-IDF, BM25, popular encoders, and, on subject classification only, zero-shot decoders; each method appears only on tasks that support it. Subject classification reaches 0.8887, while genre-form classification reaches only 0.2605, and several pair and clustering tasks remain near chance. Sparse methods remain competitive on classification, while TF-IDF is the fastest measured arm in the subject timing experiment. SHELF also varies bibliographic facets independently and can generate new, verifiably unseen documents after a model's training cutoff. Comparisons with LCSHBench and Project Gutenberg show that model rankings transfer more reliably than absolute scores, but SHELF scores do not estimate accuracy on production catalogue data. We release all source code and data under permissive licenses on GitHub and Hugging Face.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @shelf 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/shelf-a-synthetic-ha…] indexed:0 read:1min 2026-09-04 ·