A Survey on Foundations and Frontiers of Multimodal Agentic Frameworks: Techniques and Applications A new survey from arXiv (2608.20379v1) systematically examines how multimodality shapes agentic frameworks built on large multimodal models (LMMs), covering perception, reasoning, planning, memory, and action. The authors propose a modality-centric taxonomy linking architectural designs (delegated, late-fusion, early-fusion) to agent capabilities and review applications in robotics, GUI/web navigation, multimedia generation, and long-form video understanding, while analyzing performance and efficiency trade-offs. arXiv:2608.20379v1 Announce Type: new Abstract: Advances in large language models LLMs have fueled a wave of research into agency: the ability to reason, plan, and act. This effort has produced agentic frameworks that orchestrate perception, memory, and decision-making around powerful LLM backbones. With the advent of large multimodal models LMMs , these systems can process and integrate diverse modalities, including images, audio, and video, thereby improving their real-world applicability. Yet, while surveys of LLM-based agents exist, the role of multimodality in shaping agency has not been systematically examined in recent years. This survey fills the gap by analyzing the impact of multimodality across the core functional modules of the agentic framework: perception, reasoning, planning, memory, and action. Using this lens, we trace the evolution from text-centric agents to multimodal frameworks, examine how modalities are integrated through delegated, late-fusion, and early-fusion architectures, and assess the emergence of agentic behaviors enabled by grounded perception and multimodal reasoning. We organize existing work through a modality-centric taxonomy that links architectural design choices to agent capabilities. Moreover, we review multimodal agentic systems across various application domains, including Robotics, GUI & Web Navigation, Multimedia Content Generation & Editing, and Long-form Video Understanding & Retrieval. Beyond capabilities, we analyze performance across these settings and discuss efficiency-scalability trade-offs, including training and inference costs, latency, and deployment constraints. By focusing on the impact of multimodality in agentic design, we aim to identify key gaps and chart a roadmap toward robust and general-purpose intelligent systems.