17:50
2026-08-23
promptcube3.com
artificial-intelligence
Stop assuming a model is "blind" to new attacks just because the
A linear probe on frozen embeddings from Meta's Prompt Guard 2 achieves an AUC of approximately 0.999 on out-of-distribution data, revealing that the model's poor attack detection stems from a miscaliβ¦