Program Verification Makes Vision-Language Self-Evolution More Reliable
Introduction
For a vision-language model to evolve from unlabeled images, generating more questions is only half the problem. The harder challenge is assigning reliable answers to those questions. Existing pipelines often sample several model answers and use majority voting, or ask another model to act as a judge. Both approaches can turn systematic visual mistakes into training labels, allowing errors to be reinforced in later rounds.
Researchers at Mohamed Bin Zayed University of Artificial Intelligence propose VQS, or Verifiable QA Generation for Self-Evolving Models. Its central idea is to change what the model is asked to verify: instead of choosing the best complete answer, it checks a set of small, explicit visual facts. The implementation has also been released publicly.
A structured and programmatic pipeline
VQS separates image understanding, question generation, and answer computation into three stages:
- Parse the image: The model converts an image into a structured record, such as a scene graph, a chart table, or a graph representing a diagram.
- Generate and solve questions with programs: Fixed programs read the record, construct questions, and compute their answers. The answer is therefore derived from the record rather than freely guessed by a language model.
- Verify local claims: The vision-language model checks only the short facts actually read by the program, such as whether an object exists or whether a value belongs to a particular row.
These claim-level checks are also used to select training targets for the parser. As a result, the parser can improve from unlabeled images instead of relying solely on a fixed initial representation.
Why local verification matters
Majority voting assumes that the most frequent answer is probably correct. That assumption becomes weak when multiple sampled responses share the same visual bias. Model judging introduces another failure mode: the judge may be influenced by wording, reasoning style, or its own perception errors. VQS reduces the scope of the decision. The model verifies short statements, while deterministic code handles operations that are easier to execute mechanically, including combinations, relations, and calculations.
The reported human evaluation illustrates the difference. VQS answers were judged correct 94% of the time, compared with 76% for majority voting. The study also found that 24% of majority-vote labels produced during self-evolution were wrong, while model-judge labels had an 18% error rate. These findings suggest that improving the reliability of synthetic supervision can matter as much as increasing its volume.
Results and broader implications
Across ten benchmarks, VQS improved Qwen3-VL at the 2B, 4B, and 8B scales by up to 3.18 points. It outperformed the strongest self-evolving baseline at every reported size. The gains continued through three training rounds, reaching 3.84 points for the 2B model, indicating that the method is not merely a one-time filtering trick.
VQS points to a broader recipe for multimodal self-training: let the model perceive the image and propose a structured interpretation, let fixed programs perform verifiable generation and computation, and use the model to check only the visual facts that matter. This design cannot eliminate parsing errors, and automatic labels are not guaranteed to be correct. Its value is that mistakes are exposed in smaller, more inspectable units, reducing the chance that a vague end-to-end judgment becomes repeated supervision. For visual question answering with limited human annotations, program-constrained verification is a promising direction.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...