Back to articles
Evaluation & Benchmarks

Beyond Legible Letters: UltraText Bench Tests Whether Image Models Can Get Text Right

3 min read

Introduction

Generating a short, readable word inside an image is no longer the only meaningful test of visual text rendering. Real design scenes are considerably harder: a single image may contain a headline, road signs, menu items, package details, and several smaller information panels. A model must reproduce the requested strings, place them in the right locations, and preserve their visual properties at the same time. UltraText Bench, introduced by a Westlake University team, is designed to measure this denser and more sustained form of capability.

What the benchmark changes

  • Broad bilingual coverage: The benchmark includes 432 prompts from 24 real-world scene categories and three difficulty levels. English and Chinese prompts are evenly split, making language-specific behavior easier to compare.
  • Multiple text regions per image: Each human-reviewed prompt specifies exact strings for four to twelve regions. Structured references describe not only the content, but also placement and visual attributes. This turns the task from rendering one prominent word into following several simultaneous constraints.
  • Separate evaluation dimensions: Q-Judger, a vision-language model, scores each image on text fidelity, text clarity, spatial quality, and scene quality. The separation is important because an image can contain crisp-looking characters without preserving the requested wording.
  • A visible clarity–fidelity trade-off: Under the reported settings, Z-Image-Turbo improves clarity by 3.81 points over Z-Image-Base, while losing 14.76 points in fidelity. The comparison illustrates why visual sharpness alone is not a sufficient measure of text rendering.
  • Performance degrades with workload: Qwen-Image-2512’s English composite score drops from 86.50 at difficulty L1 to 42.86 at L3. Increasing the amount and complexity of text therefore reveals weaknesses that short-string tests may miss.

Why it matters

UltraText Bench is useful because it frames visual text generation as a multi-constraint execution problem rather than a simple readability check. In practical design workflows, the model must obey copy, position, hierarchy, style, and composition together. Checking only whether one headline looks plausible can significantly overestimate real-world reliability.

The benchmark also argues against relying on a single aggregate score. One model may produce sharp, highly legible glyphs while altering important words, numbers, or sentences. Another may preserve more of the requested content but struggle with layout or readability. Reporting separate dimensions gives researchers and developers more actionable evidence for model selection, data construction, and optimization.

Ten participants were involved in a human evaluation of the automatic scores, providing an additional reference for the Q-Judger-based assessment. Still, the available material describes the benchmark and selected comparisons rather than establishing a universal ranking across every language or application. Its broader message is methodological: dense, bilingual, multi-region, and progressively harder text scenes should become standard parts of image-generation evaluation.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles

CCTest · Blog
SWE-Game Tests Whether Coding Agents Can Build Playable Games
Evaluation & Benchmarks
cctest.ai

SWE-Game Tests Whether Coding Agents Can Build Playable Games

SWE-Game evaluates coding agents across 247 game-development tasks, from implementing mechanics to repairing faults and porting projects between engines. The results show that agents can produce playable prototypes, but still struggle with complete requirements and reliable gameplay logic.

Read more