MaLiang-Harness Moves Visual Generation Beyond Runnable Code
Introduction
Programmatic image and video generation promises more explicit control than a one-shot visual prompt. Yet executable code is only the beginning. A script may complete without errors and still place objects incorrectly, miss an intended appearance, or produce motion that does not match the request. MaLiang-Harness names this mismatch the Program-to-Visual (P2V) gap and frames visual generation as a persistent process rather than a single coding attempt.
Core ideas
- Persistent executable state. The framework introduces Persistent Executable Generation (PEG), which preserves the evolving visual program together with the task context. A model can therefore continue from an earlier revision instead of rebuilding the entire solution.
- Evidence linked to edits. Traceable Generation Process (TGP) connects program changes with rendered evidence. This gives the system a way to associate a visual difference with a particular edit, rather than judging only the final output.
- Revision-aware checking. Revision-aware Editing and Verification (REV) supports restoration and checks the active revision before completion. The goal is to verify that the current version—not merely an earlier successful render—meets the requested visual constraints.
- A coordinated generation loop. Planning, execution, rendering, visual feedback, and editing are organized across rendering backends within one workflow.
Evaluation findings
The authors evaluate image and video generation with MaLiang-IBench and MaLiang-VBench. The benchmarks measure generation success, visual quality, and computational cost. Eleven closed-source multimodal large language models are evaluated for image tasks, while four are evaluated for video tasks. According to the supplied paper summary, GPT-6-Astra reaches a 100% generation-success rate on both benchmarks. The share of tasks meeting all quality thresholds is 96.0% for images and 76.9% for videos.
The result is important because success rates and quality thresholds capture different parts of the problem. A model can produce runnable code consistently while still failing the visual specification. The comparison also reports a substantial mismatch between general capability scores and visual-generation performance: models with similar broad scores can differ considerably when asked to translate programs into visual outcomes.
Why it matters
MaLiang-Harness is best understood as an evaluation and orchestration framework, not simply another code-generation model. It makes the intermediate program, its history, and the rendered evidence part of the same revision reference. That design can support more systematic studies of how multimodal models diagnose visual errors, preserve useful edits, and improve outputs over multiple iterations.
The work also challenges the assumption that general benchmark performance is a sufficient proxy for visual generation ability. Future evaluations may need to test not only whether a model can write executable code, but also whether it can inspect a render, identify the source of a mismatch, and revise the correct version. The supplied material does not establish that the framework will generalize to every renderer or open model, but it offers a clear engineering pattern for making the generation-inspection-revision loop measurable and reproducible.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...