{"slug": "deberta-sentinel-toward-transparent-and-trustworthy-detection-of-ai-generated", "title": "DeBERTa-Sentinel: Toward Transparent and Trustworthy Detection of AI-Generated Text", "summary": "Researchers introduced DeBERTa-Sentinel, an AI-generated text detection framework using DeBERTa-v3's disentangled attention, achieving 98.21% validation accuracy and 97.53% test accuracy on the GLC-AIText dataset of 28,057 samples, surpassing the RoBERTa-Sentinel baseline from NeurIPS 2025. The model provides token-level explanations for transparency, supporting journalists, educators, and platform trust and safety teams in auditing detection outcomes. Code and data are available on GitHub.", "body_md": "arXiv:2608.01046v1 Announce Type: new\nAbstract: The rapid spread of large language models (LLMs) across the web raises concerns about misinformation, academic integrity, automated content manipulation, and risks to vulnerable online communities. Existing transformer-based detectors, such as GPT-Sentinel, show promise but struggle to generalize to diverse model outputs and paraphrasing attacks, limiting their role in building trustworthy web ecosystems. This work introduces DeBERTa-Sentinel, a responsible AI-generated text detection framework leveraging DeBERTa-v3's disentangled attention to capture subtle structural irregularities in synthetic content. A central design principle is transparency: unlike black-box commercial detectors, DeBERTa-Sentinel exposes token-level explanations of its decisions, enabling affected stakeholders journalists, educators, and platform trust and safety teams to audit, challenge, and contextualize detection outcomes. Using the GLC-AIText dataset of 28,057 human and LLM-generated samples (GPT, LLaMA, and Claude) with a 60-20-20 split, DeBERTa-Sentinel achieves 98.21\\% validation accuracy and surpasses the RoBERTa-Sentinel baseline from NeurIPS 2025, achieving 97.53\\% test accuracy, 95.89\\% precision, 99.33\\% recall, and 99.53\\% ROC-AUC, and maintaining a 0.665\\% false negative rate. The model's interpretability reveals linguistic markers such as academic phrasing and formal transitions associated with synthetic text, directly supporting stakeholder needs for verifiable, auditable content-authenticity decisions. By advancing responsible detection methods that reduce bias and enhance explainability, DeBERTa-Sentinel promotes trustworthy, ethical, and human-centric AI systems. Code and data are available at https://github.com/Galileo-Galili/HUMAN-VS-AI-TEXT-DETECTION.", "url": "https://wpnews.pro/news/deberta-sentinel-toward-transparent-and-trustworthy-detection-of-ai-generated", "canonical_source": "https://www.machinebrief.com/news/deberta-sentinel-toward-transparent-and-trustworthy-detectio-ujyh", "published_at": "2026-08-04 04:00:00+00:00", "updated_at": "2026-08-04 07:33:22.193751+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-safety", "ai-ethics", "ai-research"], "entities": ["DeBERTa-Sentinel", "DeBERTa-v3", "GPT-Sentinel", "RoBERTa-Sentinel", "GLC-AIText", "NeurIPS 2025", "GitHub"], "alternates": {"html": "https://wpnews.pro/news/deberta-sentinel-toward-transparent-and-trustworthy-detection-of-ai-generated", "markdown": "https://wpnews.pro/news/deberta-sentinel-toward-transparent-and-trustworthy-detection-of-ai-generated.md", "text": "https://wpnews.pro/news/deberta-sentinel-toward-transparent-and-trustworthy-detection-of-ai-generated.txt", "jsonld": "https://wpnews.pro/news/deberta-sentinel-toward-transparent-and-trustworthy-detection-of-ai-generated.jsonld"}}