A Survey on Foundations and Frontiers of Multimodal Agentic Frameworks: Techniques and Applications
06:00 · August 24, 2026 · arXiv cs.AI RSS

Advances in large language models (LLMs) have fueled a wave of research into agency: the ability to reason, plan, and act. This effort has produced agentic frameworks that orchestrate perception, memory, and decision-making around powerful LLM backbones. With the advent of large multimodal models (LMMs), these systems can process and integrate diverse modalities, including images, audio, and video, thereby improving their real-world applicability. Yet, while surveys of LLM-based agents exist, the role of multimodality in shaping agency has not been systematically examined in recent years. This survey fills the gap by analyzing the impact of multimodality across the core functional modules of the agentic framework: perception, reasoning, planning, memory, and action. Using this lens, we trace the evolution from text-centric agents to multimodal frameworks, examine how modalities are integrated through delegated, late-fusion, and early-fusion architectures, and assess the emergence of agentic behaviors enabled by grounded perception and multimodal reasoning. We organize existing work through a modality-centric taxonomy that links architectural design choices to agent capabilities. Moreover, we review multimodal agentic systems across various application domains, including Robotics, GUI & Web Navigation, Multimedia Content Generation & Editing, and Long-form Video Understanding & Retrieval. Beyond capabilities, we analyze performance across these settings and discuss efficiency-scalability trade-offs, including training and inference costs, latency, and deployment constraints. By focusing on the impact of multimodality in agentic design, we aim to identify key gaps and chart a roadmap toward robust and general-purpose intelligent systems.
Summary
Advances in large language models have driven extensive work on agency, defined as the capacity to reason, plan, and act within structured frameworks. These agentic systems coordinate perception, memory, and decision-making around LLM backbones. The arrival of large multimodal models extends this capability by enabling the processing and integration of images, audio, video, and other data streams, which in turn supports more direct interaction with physical and digital environments.
Existing surveys have addressed LLM-based agents, yet none have systematically examined how multimodality reshapes the underlying mechanisms of agency. This survey addresses that omission by tracing the shift from text-centric designs to multimodal ones and by evaluating multimodality’s effects on the five core functional modules: perception, reasoning, planning, memory, and action. It distinguishes three principal integration strategies—delegated processing, late fusion, and early fusion—and shows how each influences the emergence of grounded perception and multimodal reasoning.
A modality-centric taxonomy organizes the literature by linking specific architectural choices to observable agent capabilities. The survey then maps these systems onto application domains that include robotics, graphical-user-interface and web navigation, multimedia content generation and editing, and long-form video understanding and retrieval. Across these settings it compares empirical performance while weighing efficiency and scalability considerations such as training and inference costs, latency, and deployment constraints.
By concentrating on the role of multimodality in agent design, the work identifies persistent gaps and outlines a path toward more robust, general-purpose intelligent systems.
Why it matters
This survey is highly relevant for Dutch AI researchers and R&D teams, particularly those in the high-tech and robotics sectors, as it provides a structured taxonomy of state-of-the-art multimodal agents. It offers actionable insights into architectural choices and scalability trade-offs crucial for developing robust, real-world AI systems in the Netherlands.










