{"slug": "benchmarking-classical-and-transformer-based-models-for-document-sensitivity", "title": "Benchmarking Classical and Transformer-Based Models for Document Sensitivity Classification", "summary": "A new arXiv paper introduces Strategic 16K, a leakage-controlled corpus of 16,000 diplomatic cables from the WikiLeaks Public Library of US Diplomacy (PlusD), and benchmarks six model architectures for automatic document sensitivity classification. On the clean benchmark, BERT achieves the highest accuracy (89.14%) and F1 score (89.33%), followed by ELECTRA (Accuracy = 88.57%, F1 = 88.90%), while TF-IDF with Logistic Regression leads classical models at lower computational cost. The study addresses label leakage by removing three categories of residual classification markers, providing the first fully reproducible sensitivity classification benchmark under explicit leakage-controlled conditions.", "body_md": "arXiv:2608.16928v1 Announce Type: new\nAbstract: Automatic sensitivity classification of organizational documents is a critical yet underserved problem, where the consequences of misclassification range from regulatory violations to security breaches. While AI-based approaches offer a scalable alternative to manual review, their reliability depends fundamentally on the integrity of training data. A pervasive but underreported problem in this domain is label leakage: residual classification markers embedded within document bodies that allow models to exploit surface shortcuts rather than learning genuine content-based sensitivity signals, producing performance estimates that are inflated and unreliable. This paper addresses this problem by introducing Strategic 16K, a carefully constructed, leakage-controlled corpus of 16,000 diplomatic cables sourced from the WikiLeaks Public Library of US Diplomacy (PlusD), and presents a systematic benchmark evaluating six model architectures spanning classical machine learning and transformer-based approaches. We document an extended leakage removal protocol that identifies and eliminates three categories of residual classification markers embedded within document bodies. On the clean benchmark, BERT achieves the strongest performance (Accuracy = 89.14%, F1 = 89.33%), followed by ELECTRA (Accuracy = 88.57%, F1 = 88.90%). Among classical models, TF-IDF with Logistic Regression achieves the strongest performance at significantly lower computational cost. These results constitute the first fully reproducible sensitivity classification benchmark constructed under explicit leakage-controlled conditions from WikiLeaks PlusD.", "url": "https://wpnews.pro/news/benchmarking-classical-and-transformer-based-models-for-document-sensitivity", "canonical_source": "https://arxiv.org/abs/2608.16928", "published_at": "2026-08-19 04:00:00+00:00", "updated_at": "2026-08-19 04:13:23.774320+00:00", "lang": "en", "topics": ["machine-learning", "natural-language-processing", "artificial-intelligence"], "entities": ["arXiv", "WikiLeaks", "BERT", "ELECTRA", "TF-IDF", "Logistic Regression", "Strategic 16K", "PlusD"], "alternates": {"html": "https://wpnews.pro/news/benchmarking-classical-and-transformer-based-models-for-document-sensitivity", "markdown": "https://wpnews.pro/news/benchmarking-classical-and-transformer-based-models-for-document-sensitivity.md", "text": "https://wpnews.pro/news/benchmarking-classical-and-transformer-based-models-for-document-sensitivity.txt", "jsonld": "https://wpnews.pro/news/benchmarking-classical-and-transformer-based-models-for-document-sensitivity.jsonld"}}