{"slug": "automated-pipeline-for-herbarium-label-digitization", "title": "Automated pipeline for herbarium label digitization", "summary": "Researchers introduced HERBIOME, a modular pipeline for automated herbarium label digitization that integrates YOLOv8-based detection, CRAFT Hezar text localization, fine-tuned TrOCR, and GPT-4o Mini, achieving a Character Error Rate of 4.05-4.10% on mixed handwritten and printed text. In end-to-end tests on 450 French specimens, the pipeline reached Maximum Window Similarity of 0.614-0.618 and Semantic Metadata Accuracy of 0.440-0.445, with taxonomic fields as the main bottleneck. The work aims to unlock metadata from over 100 million specimen images for biodiversity research and multimodal AI.", "body_md": "arXiv:2608.28676v1 Announce Type: new\nAbstract: Digitized herbarium collections, now comprising over 100 million freely accessible specimen images, have become a critical resource for addressing fundamental questions in ecology and evolutionary biology. Yet the rich metadata encoded in herbarium labels (collector identities, geographic localities, collection dates, and ecological observations) remains largely inaccessible at scale, constraining both biodiversity informatics and the construction of specimen-specific image-text corpora for multimodal AI. We present HERBIOME, a modular end-to-end pipeline for automated herbarium label digitization, integrating YOLOv8-based component detection, CRAFT Hezar word-level text localization, fine-tuned TrOCR for recognition of mixed handwritten and printed text, and GPT-4o Mini for semantic metadata structuring into standardized fields. TrOCR was trained on a multi-source dataset combining general transcription corpora (CREMMA-AN, PictoCatalogs) with herbarium-specific data (R\\'eColNat), achieving a Character Error Rate of 4.05-4.10%. End-to-end evaluation on 450 French herbarium specimens, using a dual-metric framework of Maximum Window Similarity (MWS: 0.614-0.618) and Semantic Metadata Accuracy (SMA: 0.440-0.445), reveals that hybrid training strategies improve semantic fidelity while random sampling maximizes surface similarity, with taxonomic fields remaining the principal bottleneck. By automating the extraction of structured metadata from complex, heterogeneous labels, HERBIOME reduces transcription burden, enables the construction of paired image-text datasets that faithfully capture specimen individuality, which is a prerequisite for next-generation multimodal biodiversity AI systems.", "url": "https://wpnews.pro/news/automated-pipeline-for-herbarium-label-digitization", "canonical_source": "https://arxiv.org/abs/2608.28676", "published_at": "2026-09-01 04:00:00+00:00", "updated_at": "2026-09-01 04:22:28.545659+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "computer-vision", "natural-language-processing"], "entities": ["HERBIOME", "YOLOv8", "CRAFT Hezar", "TrOCR", "GPT-4o Mini", "CREMMA-AN", "PictoCatalogs", "RéColNat"], "alternates": {"html": "https://wpnews.pro/news/automated-pipeline-for-herbarium-label-digitization", "markdown": "https://wpnews.pro/news/automated-pipeline-for-herbarium-label-digitization.md", "text": "https://wpnews.pro/news/automated-pipeline-for-herbarium-label-digitization.txt", "jsonld": "https://wpnews.pro/news/automated-pipeline-for-herbarium-label-digitization.jsonld"}}