{"slug": "the-latent-diagnostic-taxonomy-a-framework-for-constructing-classifiers-and-to", "title": "The Latent Diagnostic Taxonomy: A Framework for Constructing Classifiers and Diagnosing Their Decisions, Applied to Prompt Injection Detection", "summary": "A new framework, the Latent Diagnostic Taxonomy, enables construction of classifiers and diagnosis of their decisions, applied to prompt injection detection. The framework identifies that approximately 77% of confident decisions are not robust to removing a single token, separating into two failure patterns: confidence calibration failure and exploitable shortcut. The taxonomy provides guidelines for flagging prompts that require different treatments, including relying safely, flagging heuristic bias/override, and routing insufficient context cases for human review.", "body_md": "arXiv:2608.26423v1 Announce Type: new\nAbstract: This paper proposes a framework for constructing a classifier as a safeguard layer, and for developing a complementary diagnostic that identifies which of the classifier's confident decisions can be trusted. This framework, the Latent Diagnostic Taxonomy, consists of (i) constructing a dimensionality-optimized classifier, in which the embedding dimensionality is empirically selected via cross-validated performance rather than fixed a priori, (ii) locating a relatively small set of latent support vectors (~ 29% of total training examples) representing influential prompts for identifying tokens that alter the classifier's predicted labels, and (iii) utilizing such tokens and their associated attack magnitudes for constructing a diagnostic taxonomy. This diagnostic taxonomy provides an end-to-end guideline for flagging prompts that require different treatments: rely Safely on the classifier's decision; flag Heuristic Bias and Heuristic Override cases; route Insufficient Context cases for further human/safety review. Applying the framework to a classifier trained on a public prompt injection dataset, we find that a substantial fraction of its confident decisions (~ 77%) are not robust to removing a single token, and that this brittleness separates into two distinct failure patterns: a confidence calibration failure and a genuinely exploitable shortcut. For each zone of the taxonomy, we also recommend strategies for remediating diagnosed prompts. We illustrate the framework as a series of steps, demonstrating how each step operates.", "url": "https://wpnews.pro/news/the-latent-diagnostic-taxonomy-a-framework-for-constructing-classifiers-and-to", "canonical_source": "https://arxiv.org/abs/2608.26423", "published_at": "2026-08-28 04:00:00+00:00", "updated_at": "2026-08-28 04:21:06.976374+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-safety", "ai-research"], "entities": ["Latent Diagnostic Taxonomy"], "alternates": {"html": "https://wpnews.pro/news/the-latent-diagnostic-taxonomy-a-framework-for-constructing-classifiers-and-to", "markdown": "https://wpnews.pro/news/the-latent-diagnostic-taxonomy-a-framework-for-constructing-classifiers-and-to.md", "text": "https://wpnews.pro/news/the-latent-diagnostic-taxonomy-a-framework-for-constructing-classifiers-and-to.txt", "jsonld": "https://wpnews.pro/news/the-latent-diagnostic-taxonomy-a-framework-for-constructing-classifiers-and-to.jsonld"}}