Do Audio LLMs Listen Before Acting? VGBench Tests Context Gating
Introduction
A voice assistant’s most consequential mistake may not be misunderstanding speech. It may be understanding the words correctly but failing to recognize that the words were not directed at the assistant. In a room with multiple speakers, a person can talk to someone else, think aloud, or repeat a phrase from a distance. A capable agent must decide whether it is being addressed before it answers or invokes a tool.
The paper “Do Audio LLMs Listen Before They Act?” introduces VGBench to examine this decision at the action level.
Key findings
- The benchmark evaluates action, not only recognition. VGBench contains 1,018 diagnostic items and uses a shared action space: remain silent, call a tool, or provide a natural-language answer. This makes it possible to test whether a model converts an acoustic interpretation into an appropriate action.
- The scenarios target common false activations. The suite covers side-talk, self-talk, and speaker-switch conditions. In the controlled switch pairs, the specified words remain fixed while the source, distance rendering, and temporal boundary change from a wearer-to-bystander setting. The design therefore probes whether models use contextual acoustic cues rather than just matching words.
- Raw models struggle to withhold action. Six raw Audio LLMs and three training-free adaptations often identify the intended tool, but rarely choose silence after the speaker changes. The highest raw switch mute rate is only 14%.
- Post-training substantially improves gating. In a VoxGate case study, supervised training muted 91.3% of switched commands while selecting the correct tool for all nearby wearer commands and text-only controls. An exploratory GRPO stage delivered similar switch behavior, increased side-talk accuracy from 68.4% to 70.9%, and raised self-talk muting from 52.0% to 60.0%.
Why it matters
VGBench turns “was the assistant being addressed?” from an informal product concern into a measurable diagnostic problem. Its results suggest that action gating is not equivalent to isolated speaker identification. A reliable agent must combine source changes, distance cues, temporal boundaries, and conversational context before deciding what to do.
Factorized controls identify an independent effect from changing the sound source, while sensitivity to the far-field manipulation varies across acoustic renderings. This variation matters for real deployments: a model that succeeds under one simulated recording may still react differently under another rendering of the same spatial relationship.
For smart speakers, earbuds, and tool-using voice agents, silence is a capability, not an absence of capability. Future evaluations should therefore ask not only whether a model recognized the command or selected the correct tool, but also whether it knew the command was meant for it—and whether it could refrain from acting when that judgment was uncertain.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...