Academic libraries face ongoing difficulty in aligning their collections with the needs of the communities they serve. Although the literature identified the value of using research, publications, and teaching data to inform collection development, maintaining crosswalks between these activities and library classification schemas made such work at scale impractical. This study developed a scalable workflow that maps institutional outputs to Library of Congress (LC) classification categories using a large language model to support an ongoing, evidence-based approval plan review. The workflow integrated six datasets across two domains of academic activity: research outputs and teaching activities. Where bibliographic metadata existed, they were used directly; otherwise, Anthropic’s Claude Sonnet inferred LC classifications through data-specific prompts. A Python pipeline combining the Claude and bibliographic metadata APIs produced enriched datasets at scale, followed by automated LC range validation and selective human review. The outputs feed a three-page interactive Tableau dashboard: a summary landing page, a Research Outputs Explorer, and a Teaching Activities Explorer. The collection development team has used the dashboard alongside expenditure and usage data to review approval plans, highlighting the value of unifying previously siloed data sources. The workflow is designed to refresh with new outputs and improved models, offering a sustainable foundation for ongoing collection assessment, including gap analysis and alignment with institutional priorities.
Exploring AI-assisted Classification for Collection Analysis
A study developed a scalable workflow that uses Anthropic's Claude Sonnet large language model to map institutional research and teaching outputs to Library of Congress classification categories, integrating six datasets across two domains of academic activity. A Python pipeline combining the Claude and bibliographic metadata APIs produced enriched datasets at scale, followed by automated LC range validation and selective human review, feeding a three-page interactive Tableau dashboard. The collection development team has used the dashboard alongside expenditure and usage data to review approval plans, and the workflow is designed to refresh with new outputs and improved models.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.