Omni-Perception Policy Optimization for Multimodal Emotion Reasoning
06:00 · June 25, 2026 · arXiv cs.AI RSS

We find that current emotion-oriented Omni-MLLMs still lack reliable omni-modal perception: they (i) underutilize multimodal cues in their reasoning trajectories and (ii) exhibit unfaithful behavior, often hallucinating modality-specific statements from other modalities. Building on these insights, we propose OPPO (Omni-Perception Policy Optimization), a reinforcement learning framework that explicitly optimizes multimodal perception. First, an Omni-Perception Reward decomposes ground-truth reasoning into fine-grained visual, acoustic, and emotion cues and rewards trajectories that semantically recover these cues. Second, an Omni-Perception Loss compares the policy under full and unimodally masked inputs, applying a KL penalty only to modality-specific evidence tokens to suppress cross-modal hallucination. We further introduce MEP-Bench, a diagnostic benchmark that quantifies utilization and faithfulness. Experiments show that OPPO achieves state-of-the-art performance on MER-UniBench and MME-Emotion, while substantially improving utilization and faithfulness scores on MEP-Bench, highlighting the importance of sufficient and faithful omni perception for multimodal emotion reasoning.
Summary
Current emotion-oriented Omni-MLLMs fall short in reliable omni-modal perception. They tend to underuse fine-grained visual and acoustic cues during reasoning and often produce unfaithful statements by hallucinating modality-specific details from other inputs, such as inferring visual features from audio alone. These shortcomings reduce both the interpretability and reliability of multimodal emotion reasoning outputs.
To address the issues, the authors introduce Omni-Perception Policy Optimization (OPPO), a reinforcement learning framework that directly optimizes perception quality. An Omni-Perception Reward decomposes ground-truth reasoning into discrete visual, acoustic, and emotion cues, then scores generated trajectories according to how well they semantically recover those cues. Complementing this, an Omni-Perception Loss applies a targeted KL penalty: it compares model behavior on full versus unimodally masked inputs and penalizes only the tokens that describe modality-specific evidence, thereby discouraging cross-modal hallucination while preserving overall generation stability.
The work also presents MEP-Bench, a diagnostic benchmark that measures both utilization through recall of human-annotated multimodal cues and faithfulness through POPE-style probes under unimodal masking. Experiments show that OPPO reaches state-of-the-art results on MER-UniBench and MME-Emotion while markedly raising utilization and faithfulness scores on MEP-Bench, confirming that explicit optimization of grounded, modality-faithful perception improves multimodal emotion reasoning.
Why it matters
This research is highly relevant for Dutch AI researchers focusing on trustworthy and transparent AI, as it provides novel methods to reduce hallucinations and improve the faithfulness of multimodal models. The introduction of a new benchmark and RL framework offers actionable tools for advanced practitioners developing reliable emotion-oriented AI systems.





