WorldAuditBench Tests Whether Multimodal Agents Can Audit 3D Worlds
Introduction
A capable agent in a virtual world must do more than describe what is in front of its camera. It must decide where to move, which area deserves closer inspection, and whether another observation is needed before making a judgment. WorldAuditBench studies this combined ability by asking multimodal agents to audit interactive 3D environments for defects and inconsistencies.
What the benchmark measures
The benchmark includes 213 anomaly tasks distributed across 13 interactive environments created with Unreal Engine 5 and Three.js. The tasks cover five anomaly families, with examples such as floating objects, walls that can be traversed, and objects that do not fit their surrounding scene. An agent cannot solve the task reliably from a single frame: it has to navigate, inspect, gather evidence, and report the anomaly under a fixed exploration budget.
The task therefore combines two capabilities:
- Action and search: planning movement, choosing an inspection order, and covering useful parts of the environment efficiently.
- Visual reasoning and validation: spotting suspicious evidence and then using a new viewpoint, closer inspection, or further movement to determine whether the suspicion is justified.
Two auditing paradigms
The study evaluates two ways to couple these capabilities. In the first, a vision-language-action model explores the environment, after which a vision-language model identifies the anomaly. This separates navigation from interpretation. In the second, an end-to-end VLM agent uses its visual reasoning directly to choose the next action, creating a tighter loop between observation, movement, and verification.
Across five frontier models and both paradigms, success rates range from 6.6% to 42.3%. Human performance reaches 83.4%. The gap suggests that strong image description alone is not enough. A model may recognize a scene while still failing to search systematically, revisit an uncertain location, or collect enough evidence to support its final answer.
Why it matters
WorldAuditBench is valuable because it distinguishes between seeing an anomaly and finding one. It evaluates not only recognition, but also exploration efficiency, evidence gathering, and the ability to connect an observation with a targeted action. That makes it relevant to embodied agents, world models, and interactive AI systems that must operate rather than merely answer questions about images.
The results point to a broader research challenge. Future systems will need exploration policies that account for uncertainty, actions chosen to test specific hypotheses, and reasoning processes that remain grounded in accumulated observations. Improving visual perception alone is unlikely to close the gap. The central problem is the coupling of action and reasoning in a dynamic 3D environment.
Comments
Checking sign-in status...
Loading comments...