Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs A survey examines inference-efficiency mechanisms in video and audiovisual large language models (VideoLLMs), which pair video representations with pretrained large language models and condition generation on a textual prompt. The survey attributes the high cost of video processing to these systems' strong performance across captioning, question answering, retrieval and temporal grounding tasks. No specific cost figures, dates, or benchmark numbers were provided in the available source text. Video understanding has rapidly evolved toward video large language models VideoLLMs : systems that couple video representations with pretrained large language models and condition generation on a textual prompt. Their strong performance on captioning, question answering, retrieval and temporal gro