Beyond Legible Letters: UltraText Bench Tests Whether Image Models Can Get Text Right
Introduction
Generating a short, readable word inside an image is no longer the only meaningful test of visual text rendering. Real design scenes are considerably harder: a single image may contain a headline, road signs, menu items, package details, and several smaller information panels. A model must reproduce the requested strings, place them in the right locations, and preserve their visual properties at the same time. UltraText Bench, introduced by a Westlake University team, is designed to measure this denser and more sustained form of capability.
What the benchmark changes
- Broad bilingual coverage: The benchmark includes 432 prompts from 24 real-world scene categories and three difficulty levels. English and Chinese prompts are evenly split, making language-specific behavior easier to compare.
- Multiple text regions per image: Each human-reviewed prompt specifies exact strings for four to twelve regions. Structured references describe not only the content, but also placement and visual attributes. This turns the task from rendering one prominent word into following several simultaneous constraints.
- Separate evaluation dimensions: Q-Judger, a vision-language model, scores each image on text fidelity, text clarity, spatial quality, and scene quality. The separation is important because an image can contain crisp-looking characters without preserving the requested wording.
- A visible clarity–fidelity trade-off: Under the reported settings, Z-Image-Turbo improves clarity by 3.81 points over Z-Image-Base, while losing 14.76 points in fidelity. The comparison illustrates why visual sharpness alone is not a sufficient measure of text rendering.
- Performance degrades with workload: Qwen-Image-2512’s English composite score drops from 86.50 at difficulty L1 to 42.86 at L3. Increasing the amount and complexity of text therefore reveals weaknesses that short-string tests may miss.
Why it matters
UltraText Bench is useful because it frames visual text generation as a multi-constraint execution problem rather than a simple readability check. In practical design workflows, the model must obey copy, position, hierarchy, style, and composition together. Checking only whether one headline looks plausible can significantly overestimate real-world reliability.
The benchmark also argues against relying on a single aggregate score. One model may produce sharp, highly legible glyphs while altering important words, numbers, or sentences. Another may preserve more of the requested content but struggle with layout or readability. Reporting separate dimensions gives researchers and developers more actionable evidence for model selection, data construction, and optimization.
Ten participants were involved in a human evaluation of the automatic scores, providing an additional reference for the Q-Judger-based assessment. Still, the available material describes the benchmark and selected comparisons rather than establishing a universal ranking across every language or application. Its broader message is methodological: dense, bilingual, multi-region, and progressively harder text scenes should become standard parts of image-generation evaluation.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...