{"slug": "frenchnews-7-benchmarking-cross-publisher-french-news-editorial-desk", "title": "FrenchNews-7: Benchmarking Cross-Publisher French News Editorial Desk Classification", "summary": "Researchers introduced FrenchNews-7, a cross-publisher benchmark for classifying French-language news articles into seven editorial desk categories, featuring a fine-tuned CamemBERT classifier that achieved an overall recall of 0.799, outperforming zero-shot LLM baselines (GPT-OSS-120B, Mistral Small 3.2, Llama-3.3-70B) on held-out publishers. The model and dataset are publicly available on Hugging Face.", "body_md": "arXiv:2608.18097v1 Announce Type: new\nAbstract: We present FrenchNews-7, a cross-publisher France-based French-language news editorial desk classification benchmark combining a large multi-outlet corpus, a URL-derived seven-class taxonomy, and a fine-tuned CamemBERT classifier. Labels are assigned via a hybrid pipeline combining publisher URL slugs with LLM annotation for structurally ambiguous cases, audited through an inter-rater study (2 humans + 2 LLMs; pairwise $\\kappa \\geq 0.766$, human--human $\\kappa = 0.806$). We evaluate lexical, multilingual, and French-specific trained classifiers under both in-distribution and held-out-publisher settings, with additional comparison against zero-shot LLM baselines (GPT-OSS-120B, Mistral Small 3.2, Llama-3.3-70B) on the held-out pool. The strongest model, CamemBERT-base on full article text, outperforms headline-only input, generalizes to unseen outlets, and exceeds all three zero-shot LLM baselines on overall recall (0.799), with the gap concentrated in the ambiguous editorial-boundary categories Economie and Societe. Cross-publisher evaluation reveals uneven boundary stability: Sport, Culture & Loisirs, and International transfer cleanly, while Economie (recall = 0.517) is close to blinded human agreement (0.55), and Societe (precision = 0.577) absorbs boundary ambiguity, both suggesting editorial conventions rather than recoverable classifier headroom. The fine-tuned CamemBERT-base model, labeled manifest, reference collection scripts, and a reliability-tier guidance table are available at https://huggingface.co/LeFrenchNewsLab/camembert-base-frenchnews7 (model) and https://huggingface.co/datasets/LeFrenchNewsLab/frenchnews-7 (dataset).", "url": "https://wpnews.pro/news/frenchnews-7-benchmarking-cross-publisher-french-news-editorial-desk", "canonical_source": "https://arxiv.org/abs/2608.18097", "published_at": "2026-08-20 04:00:00+00:00", "updated_at": "2026-08-20 04:12:53.147926+00:00", "lang": "en", "topics": ["natural-language-processing", "machine-learning", "large-language-models"], "entities": ["FrenchNews-7", "CamemBERT", "GPT-OSS-120B", "Mistral Small 3.2", "Llama-3.3-70B", "Hugging Face", "LeFrenchNewsLab"], "alternates": {"html": "https://wpnews.pro/news/frenchnews-7-benchmarking-cross-publisher-french-news-editorial-desk", "markdown": "https://wpnews.pro/news/frenchnews-7-benchmarking-cross-publisher-french-news-editorial-desk.md", "text": "https://wpnews.pro/news/frenchnews-7-benchmarking-cross-publisher-french-news-editorial-desk.txt", "jsonld": "https://wpnews.pro/news/frenchnews-7-benchmarking-cross-publisher-french-news-editorial-desk.jsonld"}}