Back to articles
Evaluation & Benchmarks

Duplex-MPE Tests Whether Full-Duplex Assistants Know When to Speak

3 min read

Introduction

In a conversation involving several people, a voice assistant faces a question that ordinary voice benchmarks often avoid: is this request actually directed at me? Speaking too soon can interrupt the group, while staying silent when help is expected makes the system appear unresponsive. For full-duplex speech models that can listen while speaking, deciding whether to participate is becoming as important as generating a fluent answer.

A team from Peking University introduces Duplex-MPE to study this problem. Rather than placing an assistant in a one-user dialogue, the benchmark embeds it in a shared conversation and measures whether it can participate selectively in continuous speech.

Key points

  • A multi-party setting. The benchmark includes 2,000 scenarios involving three or four human speakers and one assistant. Each request is paired across two versions: one explicitly addressing the assistant and another expressing the same request implicitly.
  • Continuous audio only. Systems receive conversation audio without transcripts or supplied turn boundaries. They must infer who is speaking, whether the request concerns them, and whether the current moment is appropriate for intervention.
  • Four separate capabilities. Duplex-MPE measures fresh response initiation, answer correctness, preservation of silence, and stopping once a human has resolved the request.
  • More speech is not automatically better. The authors evaluate five open-weight speech systems: MiniCPM-o 4.5, Moshi, FLM-Audio, Voila, and Freeze-Omni. MiniCPM-o 4.5 leads in three of the four scored capabilities. Other systems may respond more often while also producing inaccurate answers or failing to remain quiet.
  • A revealing reference comparison. A transcript-based Gemini 3.1 Pro reference responds 64.3 percentage points more often to explicit requests than to implicit ones. For the tested speech systems, paired tests find no statistically significant difference in response rates.

Why it matters

Many existing speech benchmarks focus on a designated user and evaluate turn-taking, interruption handling, or multi-round dialogue. Those abilities remain useful, but a shared conversation introduces a prior decision: does the assistant have the floor at all? Duplex-MPE treats silence as an active capability and adds the ability to stop after a person has already solved the problem, making the evaluation closer to meetings, family discussions, vehicles, and other shared environments.

The benchmark also challenges a tempting but weak proxy for quality: response frequency. An assistant that constantly jumps in may look responsive while creating friction. One that is excessively cautious may miss the moment when its help is needed. A practical full-duplex assistant therefore needs more than speech recognition and low-latency generation. It must combine speaker and reference understanding with response selection and real-time termination control.

The paired explicit-versus-implicit design is especially useful because it gives researchers a direct way to test whether a model understands addressee and conversational intent, rather than simply reacting to familiar words. The paper says that the data and code are being prepared for release. If adopted more broadly, benchmarks of this kind could shift the goal of speech systems from “sounding human” toward behaving like a considerate participant in a group conversation.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles