DiVeR Teaches VLA Verifiers to Focus on Critical Robot Decisions
Introduction
Vision-language-action (VLA) policies have become a promising foundation for general-purpose robot control. Yet scaling them with more demonstrations, larger models, and broader robotic data is expensive. Test-time scaling offers another path: a policy can generate several action candidates for the same observation, while a verifier selects the candidate most likely to lead to task success.
The Microsoft Research-led work behind DiVeR focuses on an overlooked question: which parts of a trajectory should a verifier learn to distinguish? Existing classification-based verifiers commonly use trajectory-level success or failure signals and treat visited states in roughly the same way. The paper argues that this is inefficient. At many states, plausible actions are nearly identical, so the verifier gains little from distinguishing them. A small number of states, however, admit substantially different actions that can send the robot toward very different outcomes.
Core idea
- Estimate criticality from action dispersion. DiVeR samples multiple candidate actions and examines how dispersed their representations are. Greater dispersion is used as a signal that the current state presents a more consequential choice.
- Reweight verifier training. Rather than distributing learning pressure uniformly over a trajectory, DiVeR gives greater emphasis to states where candidate actions differ more strongly.
- Avoid additional annotations. The approach does not require step-level success labels or new interaction with the environment. It extracts the training signal from candidate actions already produced during the process.
- Keep inference overhead negligible. Since the criticality estimate is derived from sampled action representations, the paper reports negligible additional verifier inference cost.
Results and implications
The authors evaluate DiVeR on LIBERO, RoboCasa, and a real Franka Research 3 robot. According to the paper summary, DiVeR consistently improves verifier-guided action selection and task success across these settings.
The broader contribution is a more targeted view of test-time scaling. Producing more candidates is useful only if the system can identify where their differences matter. DiVeR therefore shifts attention from uniform trajectory scoring toward selective discrimination at potentially decisive moments. In long-horizon manipulation, this could help a verifier spend its limited capacity on states involving contact, grasping, obstacle avoidance, or action transitions—without requiring those categories to be manually annotated.
The method also has clear limits. Action-representation dispersion is a proxy for decision criticality, not a direct measurement of future task value. Its usefulness depends on whether the representation space captures differences that actually affect downstream execution. Candidate quality and the reliability of the underlying VLA policy remain important as well. Even so, DiVeR presents a practical way to improve existing verifier-guided VLA inference without expanding the robot-data collection burden.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...