The Latent Diagnostic Taxonomy: A Framework for Constructing Classifiers and Diagnosing Their Decisions, Applied to Prompt Injection Detection A new framework, the Latent Diagnostic Taxonomy, enables construction of classifiers and diagnosis of their decisions, applied to prompt injection detection. The framework identifies that approximately 77% of confident decisions are not robust to removing a single token, separating into two failure patterns: confidence calibration failure and exploitable shortcut. The taxonomy provides guidelines for flagging prompts that require different treatments, including relying safely, flagging heuristic bias/override, and routing insufficient context cases for human review. arXiv:2608.26423v1 Announce Type: new Abstract: This paper proposes a framework for constructing a classifier as a safeguard layer, and for developing a complementary diagnostic that identifies which of the classifier's confident decisions can be trusted. This framework, the Latent Diagnostic Taxonomy, consists of i constructing a dimensionality-optimized classifier, in which the embedding dimensionality is empirically selected via cross-validated performance rather than fixed a priori, ii locating a relatively small set of latent support vectors ~ 29% of total training examples representing influential prompts for identifying tokens that alter the classifier's predicted labels, and iii utilizing such tokens and their associated attack magnitudes for constructing a diagnostic taxonomy. This diagnostic taxonomy provides an end-to-end guideline for flagging prompts that require different treatments: rely Safely on the classifier's decision; flag Heuristic Bias and Heuristic Override cases; route Insufficient Context cases for further human/safety review. Applying the framework to a classifier trained on a public prompt injection dataset, we find that a substantial fraction of its confident decisions ~ 77% are not robust to removing a single token, and that this brittleness separates into two distinct failure patterns: a confidence calibration failure and a genuinely exploitable shortcut. For each zone of the taxonomy, we also recommend strategies for remediating diagnosed prompts. We illustrate the framework as a series of steps, demonstrating how each step operates.