Back to articles
Speech & Audio

Voice Agents That Act Before You Finish Speaking

3 min read

Introduction

The delay of a voice assistant is not determined only by speech recognition or language-model inference. When a request requires weather data, a calendar lookup, search, or another external capability, a conventional cascaded system often waits for one stage to finish before starting the next: speech recognition completes, the language model identifies a tool call, the tool runs, and the model then prepares the answer. This arrangement makes the full tool delay visible only after the user has stopped speaking.

The paper Hiding Tool Latency in On-Device Cascaded Voice Agent through Speculative Execution explores a different scheduling strategy. Instead of waiting for the final transcript, the system uses partial ASR hypotheses to anticipate likely tool requests and starts work before the utterance is complete.

How the system works

  • Prediction from streaming ASR: A dedicated Predictor monitors the evolving intermediate transcript and estimates whether the user is forming a request that needs a tool.
  • Speculative execution and caching: When a candidate call is identified, the tool is launched in the background. Its output is cached and can later be inserted into the prompt sent to the language model, avoiding a fresh wait after the user finishes speaking.
  • Validation after self-correction: Early hypotheses are inherently uncertain. A user may revise a location, time, or other argument halfway through a sentence. The system therefore applies rule-based validation and injects only cached results that remain consistent with the recognized request.
  • A safe fallback: The language model retains the ability to call tools directly. If the prediction is wrong, the cache is unusable, or no speculative call was made, the system can fall back to the ordinary serialized pipeline. Speculation is therefore an optimization rather than a mandatory source of truth.

Results and implications

The authors evaluated the approach with live measurements from a fully implemented Android voice assistant. Median time-to-first-audio decreased from 5.79 seconds to 4.60 seconds, while the standard deviation dropped from 3.49 seconds to 2.81 seconds. The reported improvement is therefore not limited to a faster typical response; it also makes response timing more predictable.

The main idea is to overlap work that would otherwise be serialized. Tool execution begins during recognition, so part of its latency is hidden behind the user’s remaining speech. This is particularly relevant for on-device cascaded systems because it does not require replacing the entire architecture with a fully end-to-end or full-duplex model. A predictor, cache, validation layer, and existing tool interface can be added around the current components.

The trade-off is that speculative execution can create incorrect or unnecessary calls and may consume additional local or network resources. Validation and fallback reduce the consequences, but the supplied results do not establish how the method behaves across a wider range of tools, devices, correction patterns, or network conditions. Still, the study highlights an important systems lesson: lower voice-agent latency can come not only from faster models, but also from starting sufficiently confident work earlier.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles