EyeRobot 2.0: Active Gaze for Manipulation Without Wrist Cameras
Introduction
Wrist cameras are a common solution to a basic problem in robot manipulation: a fixed camera can lose sight of an object when an arm moves close to it, when the object is occluded, or when both hands interact in a confined workspace. The cost is additional hardware, calibration, wiring, and another sensor stream for the policy to interpret. EyeRobot 2.0 investigates whether a robot can recover some of that local visual access by actively moving its “eyes.”
How the system works
- Active visual fixation: Using a single stereo camera with two steerable eye viewpoints, the robot directs its gaze toward a 3D fixation point in the scene. The camera is therefore not merely recording the workspace; its viewpoint becomes part of the control loop.
- Foveal processing: Images are processed with more visual tokens allocated to the center than to the periphery. This mirrors the idea of high-acuity human vision and concentrates computation on the object or contact region currently relevant to the task.
- Hierarchical gaze control: A low-level gaze-servoing policy learns to align the viewpoints with a goal object. A higher-level target selector chooses fixation goals as the task progresses. Both components are trained with reinforcement learning on real-world data: the servoing policy uses a dense geometric reward, while the selector is co-trained with a behavioral-cloning gripper policy to discover useful fixation sequences.
- Fixation-relative actions: EyeRobot 2.0 expresses gripper information in an SE(3) frame relative to the fixation point. This can make the action distribution more compact than a representation tied directly to a global scene frame.
Evaluation and interpretation
The authors collected teleoperation data for seven real-world tasks and six simulated tasks. They report more than 1,000 physical robot trials and 1,800 simulated trials, comparing active gaze with passive stereo and with policies using an egocentric view plus wrist cameras. The central contribution is not simply a new camera arrangement. It is the coordination of attention, viewpoint control, visual token allocation, and manipulation policy in one loop.
The supplied material does not provide the individual success rates or the numerical gap between baselines. It would therefore be premature to claim that active gaze universally outperforms wrist-camera systems. What the setup does establish is a meaningful evaluation question: can learned, task-dependent viewing replace some of the local information normally supplied by a wrist sensor?
Why it matters
EyeRobot 2.0 reframes wrist-camera removal as a problem of physical attention rather than sensor subtraction. A robot can decide what deserves a close look, move its viewpoints accordingly, and express actions in coordinates centered on that observation. If the approach remains reliable under occlusion, contact-rich manipulation, and bimanual coordination, it could reduce hardware complexity while preserving useful local feedback.
The trade-off is a new dependency between perception and control. A poor fixation can degrade visual understanding, while an inefficient gaze sequence can disrupt the manipulation itself. Real-time execution, robustness to unexpected motion, and generalization to tasks not represented in the training data will therefore be important follow-up questions.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...