Back to articles
Evaluation & Benchmarks

Do Audio LLMs Listen Before Acting? VGBench Tests Context Gating

3 min read

Introduction

A voice assistant’s most consequential mistake may not be misunderstanding speech. It may be understanding the words correctly but failing to recognize that the words were not directed at the assistant. In a room with multiple speakers, a person can talk to someone else, think aloud, or repeat a phrase from a distance. A capable agent must decide whether it is being addressed before it answers or invokes a tool.

The paper “Do Audio LLMs Listen Before They Act?” introduces VGBench to examine this decision at the action level.

Key findings

  • The benchmark evaluates action, not only recognition. VGBench contains 1,018 diagnostic items and uses a shared action space: remain silent, call a tool, or provide a natural-language answer. This makes it possible to test whether a model converts an acoustic interpretation into an appropriate action.
  • The scenarios target common false activations. The suite covers side-talk, self-talk, and speaker-switch conditions. In the controlled switch pairs, the specified words remain fixed while the source, distance rendering, and temporal boundary change from a wearer-to-bystander setting. The design therefore probes whether models use contextual acoustic cues rather than just matching words.
  • Raw models struggle to withhold action. Six raw Audio LLMs and three training-free adaptations often identify the intended tool, but rarely choose silence after the speaker changes. The highest raw switch mute rate is only 14%.
  • Post-training substantially improves gating. In a VoxGate case study, supervised training muted 91.3% of switched commands while selecting the correct tool for all nearby wearer commands and text-only controls. An exploratory GRPO stage delivered similar switch behavior, increased side-talk accuracy from 68.4% to 70.9%, and raised self-talk muting from 52.0% to 60.0%.

Why it matters

VGBench turns “was the assistant being addressed?” from an informal product concern into a measurable diagnostic problem. Its results suggest that action gating is not equivalent to isolated speaker identification. A reliable agent must combine source changes, distance cues, temporal boundaries, and conversational context before deciding what to do.

Factorized controls identify an independent effect from changing the sound source, while sensitivity to the far-field manipulation varies across acoustic renderings. This variation matters for real deployments: a model that succeeds under one simulated recording may still react differently under another rendering of the same spatial relationship.

For smart speakers, earbuds, and tool-using voice agents, silence is a capability, not an absence of capability. Future evaluations should therefore ask not only whether a model recognized the command or selected the correct tool, but also whether it knew the command was meant for it—and whether it could refrain from acting when that judgment was uncertain.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles

CCTest · Blog
PhysVista Tests VLM Physical Intelligence Through a Perception–Reasoning–Assessment Loop
Evaluation & Benchmarks
cctest.ai

PhysVista Tests VLM Physical Intelligence Through a Perception–Reasoning–Assessment Loop

PhysVista introduces a benchmark that evaluates whether vision-language models understand physical consistency rather than merely recognizing visual content. It combines physical state perception, dynamics reasoning, and plausibility assessment across real-world and AI-generated videos.

Read more
CCTest · Blog
OpenTumorBoard Tests Whether AI Can Reason With a Cancer Care Team
Evaluation & Benchmarks
cctest.ai

OpenTumorBoard Tests Whether AI Can Reason With a Cancer Care Team

OpenTumorBoard turns public multidisciplinary tumor board recordings into a benchmark for evaluating models on specialist answers and full clinical discussions. Its results show that even advanced general and medical models still struggle to reproduce expert responses and board-level consensus.

Read more