{"slug": "semalith-v1-4-a-calibrated-184m-safety-classifier-achieving-state-of-the-art-at", "title": "Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B", "summary": "Semalith v1.4, a 184M-parameter DeBERTa-v3-base safety classifier, achieves state-of-the-art prompt-injection detection at 44x fewer parameters than Llama-Guard-3-8B, winning all 7 prompt-injection evaluations and 11 of 18 overall benchmarks on 22 held-out tests. Developed by the Semalith team, the model performs simultaneous three-axis safety classification (prompt injection, general harm, and financial-services regulatory compliance) in a single forward pass, with zero false positives on 208 benign agentic prompts versus 0.063 for Llama-Guard-3-8B.", "body_md": "arXiv:2607.22545v1 Announce Type: new\nAbstract: Deploying large language models in financial-services and agentic settings requires safety classifiers that simultaneously handle prompt injection, regulatory compliance, and general harm, a combination no existing open guardrail addresses in a single inference pass.\nSemalith v1.4 is a 184M-parameter DeBERTa-v3-base classifier performing simultaneous three-axis safety classification including prompt injection, general harm, and financial-services regulatory compliance, in a single forward pass. Its 22-class head (BENIGN, nine prompt-injection sub-types, general-harm, eleven BFSI labels) is trained with a 4-class auxiliary super-category head under jointly weighted loss, on a 76,204-row corpus mined from 49 public sources with SHA-1 deduplication against every held-out evaluation set, with 21 of 22 benchmarks at zero contamination (max 0.22%).\nAgainst Llama-Guard-3-8B on 22 held-out benchmarks, Semalith v1.4 wins every prompt-injection evaluation (7/7) and 11 of 18 benchmarks overall at 44x fewer parameters, with FPR = 0.000 on 208 benign agentic prompts vs 0.063 for Llama-Guard-3-8B. On general-harm benchmarks (WildGuardMix, HEx-PHI, HarmBench), Llama-Guard-3 leads; this complementary split is documented in Section 4. Six measured weak spots are disclosed in Section 6.\nDeployment guidance: v1.3 is recommended for conversational moderation deployments (ToxicChat F1 0.624); v1.4 is recommended when BFSI label coverage or zero-FPR on benign agentic prompts is the priority.", "url": "https://wpnews.pro/news/semalith-v1-4-a-calibrated-184m-safety-classifier-achieving-state-of-the-art-at", "canonical_source": "https://arxiv.org/abs/2607.22545", "published_at": "2026-07-28 04:00:00+00:00", "updated_at": "2026-07-28 04:11:57.092200+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-safety", "ai-research", "ai-products"], "entities": ["Semalith", "DeBERTa-v3-base", "Llama-Guard-3-8B", "WildGuardMix", "HEx-PHI", "HarmBench", "ToxicChat", "BFSI"], "alternates": {"html": "https://wpnews.pro/news/semalith-v1-4-a-calibrated-184m-safety-classifier-achieving-state-of-the-art-at", "markdown": "https://wpnews.pro/news/semalith-v1-4-a-calibrated-184m-safety-classifier-achieving-state-of-the-art-at.md", "text": "https://wpnews.pro/news/semalith-v1-4-a-calibrated-184m-safety-classifier-achieving-state-of-the-art-at.txt", "jsonld": "https://wpnews.pro/news/semalith-v1-4-a-calibrated-184m-safety-classifier-achieving-state-of-the-art-at.jsonld"}}