A Unified Evaluation Framework for Trustworthy Large Language Models, Agentic AI, and Multimodal Systems A new arXiv paper (2609.19524v1) proposes a unified evaluation framework for trustworthy AI that assesses large language models, agentic systems, and multimodal models across eight trustworthiness dimensions: capability, robustness, safety, fairness, transparency, governance, oversight, and efficiency. The framework maps native measurements to common performance bands with uncertainty estimates and traceable evidence, adds a meta-evaluation layer for validity, reliability, and reproducibility, and uses safety-critical overrides so aggregate scores cannot mask critical failures. The authors state that empirical validation across deployment contexts remains an essential next step. arXiv:2609.19524v1 Announce Type: new Abstract: Benchmark scores alone provide an incomplete basis for assessing the trustworthiness of modern artificial intelligence systems. Large language models LLMs , agentic systems, and multimodal models MLLMs require different forms of assessment, yet their evaluation evidence must remain interpretable for development and oversight. We propose a unified framework that connects output-level, trajectory-level, and cross-modal assessment through eight trustworthiness dimensions: capability, robustness, safety, fairness, transparency, governance, oversight, and efficiency. The framework preserves system-specific metrics while mapping native measurements to common performance bands, accompanied by uncertainty estimates and traceable evidence. A meta-evaluation layer examines the validity, reliability, and reproducibility of the evaluation itself. Multidimensional profiles expose strengths and weaknesses, while safety-critical overrides prevent aggregate scores from masking critical failures. Mappings to governance frameworks, international standards, and European Union regulatory requirements connect technical assessment with oversight needs. The framework provides a structured basis for assessing both system performance and the credibility of the evidence supporting it, with empirical validation across deployment contexts remaining an essential next step.