LEGO-Anything Turns a Single Image into an Editable 3D Scene Program
Introduction
Single-image 3D reconstruction is often presented as a problem of producing a mesh, point cloud, or convincing render. LEGO-Anything asks a more practical question: can the result remain inspectable, editable, and queryable after reconstruction? Its answer is an image-to-code framework in which a coding agent builds a scene by writing and executing Blender code instead of producing only a fixed 3D asset.
From visual input to executable code
The agent receives an image, proposes Blender code, runs it, inspects the generated scene and renderings, and then revises the program. This loop combines visual interpretation, geometric construction, and debugging. The output is therefore a scene program that can be executed again and potentially modified, inspected, or queried by downstream systems.
That representation also makes the weaknesses of agentic reconstruction easier to see. The authors identify three recurring problems:
- Weak initialization: if the first scene structure is poor, later edits have little solid foundation.
- Regressive edits: a new change can improve one visible region while damaging an earlier, more accurate part.
- Unreliable self-evaluation: the agent’s assessment of its own render does not always track geometric or appearance fidelity.
A benchmark that separates failure modes
The authors introduce LEGO-Bench, a simulator-grounded benchmark containing 208 images from 104 diverse indoor and outdoor scenes. Instead of collapsing quality into a single visual score, it separately measures whether the artifact is valid, whether visible-surface geometry is recovered, and whether the rendered appearance matches the reference. This design makes it easier to distinguish a scene that runs successfully from one that is actually faithful.
Among the evaluated agents, GPT-6-astra achieves the strongest overall results, scoring 53.4% indoors and 39.6% outdoors. The numbers also show a substantial gap between producing a valid scene artifact and recovering the scene’s geometry and appearance accurately. In other words, successful execution is a necessary milestone, not proof of reliable reconstruction.
Improving the agent’s construction loop
LEGO-Plugin is a training-free harness plugin designed to make iterative construction more controlled. On the 42-case Office subset, it improves all six evaluated models, with relative overall gains of up to 62.7%. The result suggests that workflow design matters: preserving useful state, constraining edits, and organizing feedback may be as important as adding more raw model capability.
A programmable representation for vision
LEGO-World tests whether reconstructed scenes can support downstream visual tasks. From scenes produced by GPT-6-astra, the researchers derive object detections, instance masks, and relative depth through deterministic scene queries. All three tasks show non-trivial performance, but they remain well behind specialist vision models. The scenes are therefore promising as a bridge between visual input and programmable environments, though not yet precise enough to serve as a dependable representation of natural images.
Why it matters
The project’s main contribution is conceptual as much as technical: it connects image understanding, code generation, and interactive 3D environments in one loop. Such a representation could be useful for simulation, editable content creation, and agents that need structured spatial information. Yet a single image leaves many surfaces occluded and depths ambiguous, while repeated code edits can accumulate errors. Future progress will require stronger initialization, reversible editing, and more trustworthy self-evaluation—not just more realistic renders.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...