{"slug": "aixamine-unified-black-box-evaluation-of-cross-dimensional-trade-offs-in-llm-and", "title": "aiXamine: Unified Black-Box Evaluation of Cross-Dimensional Trade-offs in LLM Safety, Security, and Privacy", "summary": "Researchers introduced aiXamine, a unified black-box platform that evaluates LLM trustworthiness across safety, security, and privacy, and found that stronger safety alignment increases over-refusal, privacy is near-orthogonal to other dimensions, and distillation can cause robustness collapse (56.9 to 2.6). The study, the largest joint safety, security, and privacy evaluation to date, tested over 120 LLMs through more than 5,000 runs across 46 tests and nine services.", "body_md": "arXiv:2608.20554v1 Announce Type: cross\nAbstract: The critical failure modes in deployed large language models (LLMs) are cross-dimensional: a model can score 99.3 in safety alignment while refusing one in three benign queries, or improve across every capability metric while losing 21 points in privacy. Existing evaluation frameworks that assess safety, security, and privacy independently cannot detect these patterns. We introduce aiXamine, a unified black-box platform that evaluates LLM trustworthiness across safety, security, and privacy as interdependent properties. aiXamine orchestrates 46 tests across nine services through an automated red-teaming pipeline, producing hierarchical risk profiles, from prompt-level diagnostics to cross-service trade-off analytics, that enable reproducible comparison of proprietary and open-weight systems under identical conditions. Applying aiXamine to over 120 LLMs through more than 5,000 test runs, we conduct the largest joint safety, security, and privacy study to date and uncover three cross-dimensional phenomena invisible to single-axis evaluation. First, safety enforcement incurs a quantifiable safety tax: stronger alignment systematically increases over-refusal, forcing providers to choose between protection and utility. Second, privacy is near-orthogonal to other trustworthiness dimensions and not captured by standard alignment. Third, we identify and formally characterize distillation-induced robustness collapse: off-policy distillation without on-policy correction causes entropy collapse, catastrophically destroying robustness (56.9$\\to$2.6) on the same base architecture. These findings, compounded by diminishing returns from scale and category-dependent safety behaviors, demonstrate that trustworthiness is inherently multi-dimensional: progress along one axis does not guarantee, and can actively undermine, progress along others, yet current alignment methods treat it as a single objective.", "url": "https://wpnews.pro/news/aixamine-unified-black-box-evaluation-of-cross-dimensional-trade-offs-in-llm-and", "canonical_source": "https://www.machinebrief.com/news/aixamine-unified-black-box-evaluation-of-cross-dimensional-t-pmay", "published_at": "2026-08-24 04:00:00+00:00", "updated_at": "2026-08-24 05:14:09.295949+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-safety", "ai-ethics", "ai-research"], "entities": ["aiXamine", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/aixamine-unified-black-box-evaluation-of-cross-dimensional-trade-offs-in-llm-and", "markdown": "https://wpnews.pro/news/aixamine-unified-black-box-evaluation-of-cross-dimensional-trade-offs-in-llm-and.md", "text": "https://wpnews.pro/news/aixamine-unified-black-box-evaluation-of-cross-dimensional-trade-offs-in-llm-and.txt", "jsonld": "https://wpnews.pro/news/aixamine-unified-black-box-evaluation-of-cross-dimensional-trade-offs-in-llm-and.jsonld"}}