KAIST Unveils AI Immune to Night and Smoke Errors
04:07 · August 3, 2026 · RSS APP - AI Primary Research

The research team. From left: Sangyun Chung (KAIST, first author of the MAD study and co-first author of the DNA study); Yong Man Ro (KAIST
Summary
Researchers at the Korea Advanced Institute of Science and Technology have introduced two techniques, Diverse Negative Attributes (DNA) and Modality-Adaptive Decoding (MAD), to reduce cross-modal hallucinations in multimodal large language models. These models combine inputs such as text, standard RGB images, thermal or depth data, and audio, yet they frequently misread sensor physics or generate nonexistent sounds when visual cues dominate. The new methods address both the underlying sensory bias toward ordinary camera images and the interference that arises when modalities are processed together.
DNA improves an MLLM’s grasp of non-RGB sensors by constructing the VS-TDX benchmark and treating common model errors as training signals. The approach lets the model internalize the physical principles of thermal, depth, and X-ray imagery, so that bright regions in a thermal frame are correctly attributed to emitted heat rather than reflected light. Because the adjustment uses only a modest dataset, it avoids the expense of full-scale retraining.
MAD operates at inference time without any parameter updates. For each query the model evaluates the relative reliability of vision and audio, then dynamically raises the weight of the more pertinent modality. This plug-in mechanism suppresses the tendency to invent audio events from visual objects, such as claiming the sound of splashing water when only a boat appears on screen. Both techniques were validated on tasks relevant to low-visibility navigation and multi-sensor inspection, and the same doctoral student served as lead author on the associated studies presented at CVPR and published in IEEE Transactions on Image Processing.
Why it matters
This primary research is highly relevant for Dutch AI researchers as it provides cost-efficient, actionable methodologies to reduce hallucinations in multimodal models. Improving AI reliability aligns strongly with the Netherlands' focus on trustworthy and ethical AI deployment in sectors like healthcare and autonomous systems.









