OmniConfess Traces Which Evidence Drives Multimodal Hallucinations
Introduction
Multimodal large language models can combine text, images, audio, and video, but access to more evidence does not automatically make their answers more reliable. When modalities disagree, or when the model attends to the wrong source, the resulting answer may still contain hallucinations. In many cases, the output shows that something went wrong without revealing which evidence channel encouraged the incorrect claim.
OmniConfess addresses this diagnostic problem with a training-free inference-time procedure. Rather than immediately generating a new answer or suppressing uncertain content wholesale, it examines the evidential dependence of a fixed candidate response at token resolution.
How the method works
The approach can be summarized in three stages:
- Anchor a candidate response. The system starts with a generated or supplied answer and treats it as the object to be examined.
- Intervene on evidence channels. It controls individual channels, such as text, image, audio, or video evidence, and re-scores the candidate at the token level. Changes in token scores indicate how strongly particular parts of the answer depend on each channel.
- Use the resulting confession for correction. The token-by-channel signals are assembled into a structured “confession.” Content supported by relevant evidence can be retained, while commitments driven by irrelevant or contradictory evidence can be revised.
The term “confession” refers to an evidence-dependence record, not necessarily a natural-language explanation generated by the model. Its purpose is to make the hidden support behind a response more inspectable and actionable.
Evaluation and limitations
The authors introduce OmniHalluBench, a benchmark with 3,540 examples assembled from six datasets. It covers text, image, audio, and video settings, as well as both judgment and free-form generation tasks. According to the supplied paper summary, experiments show that OmniConfess reduces hallucinations across heterogeneous modalities and task settings.
The available material does not include detailed baselines, metrics, or per-task improvements. It therefore supports the conclusion that the method is promising, but not a precise assessment of how much it improves over competing approaches. Its computational overhead, the stability of channel interventions, and its ability to assign blame when several modalities jointly cause an error also require closer examination in the full paper and implementation.
Why it matters
The main contribution is a shift from asking only whether an answer is wrong to asking what evidence sustained the wrong commitment. This is a useful direction for multimodal safety because it could enable selective correction: grounded statements need not be discarded merely because another part of the answer is unreliable.
If the approach generalizes across models and deployment settings, token-level evidence tracing could support more targeted factual correction and provide researchers with a tool for studying cross-modal conflicts. At the same time, attribution signals should not automatically be treated as complete causal explanations. Their practical reliability will depend on broader evaluation and reproducible implementation.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...