{"slug": "croissantminer-automated-extraction-and-validation-of-croissant-metadata-for-ml", "title": "CroissantMiner: Automated Extraction and Validation of Croissant Metadata for ML Datasets", "summary": "Researchers released CroissantMiner, the first benchmark for end-to-end evaluation of Croissant metadata extraction from ML dataset documentation, comprising 602 papers — 102 with human-validated gold annotations and 500 with LLM-generated silver annotations — covering the full Croissant schema including core and Responsible AI (RAI) fields. Evaluating frontier models, open-weight models, and agentic architectures under a two-tier framework combining rule-based scoring with a human-audited LLM judge, the team found single-pass extraction consistently outperformed all four agentic architectures, with the largest gap on long-form RAI fields. The benchmark, evaluation code, judge audit, live demo, and leaderboard are released openly to new systems.", "body_md": "arXiv:2610.07132v1 Announce Type: new \nAbstract: Croissant has emerged as a standard for machine-readable dataset metadata, yet populating its fields remains labor-intensive and requires careful reading of accompanying dataset documentation. We present the first benchmark enabling end-to-end evaluation of metadata extraction aligned with a community-standard schema. The benchmark comprises 602 papers, including 102 with human-validated gold annotations and 500 with LLM-generated silver annotations, covering the full Croissant schema with both core and Responsible AI (RAI) fields. Using this benchmark, we evaluate a range of extraction systems spanning frontier models, open-weight models, and agentic architectures, under a two-tier evaluation framework that combines rule-based scoring with an LLM judge selected via human audit. We find that single-pass extraction consistently outperforms the four agentic architectures we evaluate: across backbones, these decomposed variants achieve lower accuracy than a single full-context pass. The largest gap appears on long-form RAI fields, which require synthesizing and interpreting information scattered across a paper rather than copying it from a single location, a setting where current systems remain far from reliable. We release the benchmark, evaluation code, judge audit, a live demo, and a leaderboard open to new systems.", "url": "https://wpnews.pro/news/croissantminer-automated-extraction-and-validation-of-croissant-metadata-for-ml", "canonical_source": "https://arxiv.org/abs/2610.07132", "published_at": "2026-10-07 04:00:00+00:00", "updated_at": "2026-10-07 04:19:46.892880+00:00", "lang": "en", "topics": ["machine-learning", "structured-data", "ai-research", "large-language-models", "ai-agents"], "entities": ["CroissantMiner", "Croissant", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/croissantminer-automated-extraction-and-validation-of-croissant-metadata-for-ml", "markdown": "https://wpnews.pro/news/croissantminer-automated-extraction-and-validation-of-croissant-metadata-for-ml.md", "text": "https://wpnews.pro/news/croissantminer-automated-extraction-and-validation-of-croissant-metadata-for-ml.txt", "jsonld": "https://wpnews.pro/news/croissantminer-automated-extraction-and-validation-of-croissant-metadata-for-ml.jsonld"}}