VLM Judges Can Change Their Answers Even When Images Should Be Ignored
Introduction
Vision-language models are increasingly used as annotators, reviewers, and substitutes for human judgment. That makes evaluation design especially important. If a task says that the answer must be determined from text alone, an accompanying image should not matter. Yet a new study from a Technion team suggests that an image can still alter a model’s decision even when it contributes no valid evidence.
The MIST test
The researchers introduce MIST, the Misleading-Image Stress Test. It contains 200 English sentences built around phrases that can be interpreted literally or figuratively. Each sentence is evaluated under three conditions:
- an aligned image showing one interpretation;
- a misleading image showing the opposite interpretation;
- no image at all.
The evaluation instructions require the label to be based on the sentence alone. Therefore, a robust judge should return the same answer in all three settings. The study tested 13 VLM judges and compared their outputs with human annotations.
Presence matters more than content
The results do not show a consistent tendency to follow the visual evidence. Aligned images changed 20.5% of labels, while misleading images changed 19.4%. The rates were close for every judge and both were higher than the 11.6% change observed when the instruction to ignore the image was removed while the image itself remained.
The more revealing result concerns the direction of change. Among labels that differed between the aligned and misleading image conditions, only 37% moved toward the meaning depicted by the image. Thus, the models were not simply choosing the interpretation suggested by the visual input. The mere presence of an image appeared to be a stronger source of instability than the image’s semantic content.
Agreement with human annotators did not improve when an image was aligned with the sentence, nor did it collapse in a systematic way when the image was misleading. The study also separated seven judges that passed an additional alt-test from six that did not. The effect was smaller among the passing judges, but it remained present in both groups.
Why this matters for evaluation
The findings challenge a common assumption in model substitution studies: that a reported verdict primarily reflects the model’s capability. In practice, it may also reflect the prompt, modality mix, image placement, and instructions used during the test. A single evaluation configuration can therefore hide or amplify behavioral instability.
More robust testing should compare missing, aligned, conflicting, and irrelevant modalities. It should report not only accuracy, but also whether the model relies on evidence that the task actually permits. This is particularly important when VLMs are used for annotation, moderation, or quality assessment, where an unexplained label change can affect downstream data.
MIST does not show that VLMs always obey images or always ignore text instructions. Its narrower and more useful lesson is that an irrelevant modality can disturb a judgment without supplying a reliable alternative signal. A substitutability result therefore describes a complete evaluation setup—not just the model under test.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...