ArtiMo: Agent-Driven Articulated Mesh Animation
Researchers propose ArtiMo, a zero-shot agent-driven framework that uses large language and vision-language models to animate articulated 3D meshes from text without fine-tuning, leveraging URDF kinem…
Researchers propose ArtiMo, a zero-shot agent-driven framework that uses large language and vision-language models to animate articulated 3D meshes from text without fine-tuning, leveraging URDF kinem…
Researchers introduced KANEx, the first framework leveraging Kolmogorov-Arnold Networks (KANs) to ground Vision-Language Model (VLM) reasoning for medical explainability, achieving a 10% improvement i…
A new study evaluating Vision-Language Models (VLMs) for safety reasoning finds that these models frequently misinterpret anomalous scenes as hazardous, revealing an over-reliance on contextual irregu…
Researchers introduced RTSGameBench, a benchmark built on the real-time strategy game Beyond All Reason, to evaluate strategic reasoning in Vision-Language Models. The benchmark includes diverse match…
Researchers introduced GridVQA-X, a diagnostic framework to evaluate cross-modal explainability in Vision-Language Models. The framework uses synthetic data with ground-truth explanations to test whet…
Researchers have developed CoReVAD, a training-free video anomaly detection framework that uses a single frozen Vision-Language Model to generate both anomaly scores and temporal descriptions without …
Researchers have introduced Transcoders, a function-centric framework that decomposes vision-language models into interpretable computational pathways linking image patches to token generation. Applie…