SceneActBench Tests Whether VLM Agents Can Act in the 3D Scenes They See
SceneActBench shifts evaluation from describing 3D scenes to acting within them. The benchmark measures whether vision-language model agents can turn visual input into reliable actions across multi-object 3D environments.
Read more