Skill2Real Brings Executable Robot Skills from Simulation to Reality
Moving a robot policy from simulation to the physical world is difficult because perception, dynamics, and robot embodiment rarely match perfectly. A behavior that succeeds in a simulator may therefore fail when the same instructions meet camera noise, different contact dynamics, or another hardware configuration. Skill2Real proposes to address this gap by transferring executable skills and their memories instead of retraining a complete task policy for every deployment setting.
How the framework works
The central abstraction is a shared robot application programming interface, or API. Skills are expressed through public observations and API semantics rather than hidden simulator state. This gives the learned behaviors a boundary that can, in principle, remain meaningful across environments and robot embodiments.
Training is organized around a Proposer–Verifier–Governor, or PVG, loop. The Proposer creates or revises a skill. The Verifier uses privileged evidence available in simulation to inspect the outcome and diagnose why an attempt succeeded or failed. The Governor then decides how proposed changes should be accepted and how the skill memory should evolve. This design turns simulator feedback into a structured review process rather than relying only on undirected trial and error.
Skill2Real also uses a hierarchy. Its “Cerebellum” first learns local manipulation skills. A “Brain” subsequently learns task-level composition while the Cerebellum remains frozen. Both memories are then transferred to the real robot without task-policy fine-tuning or skill-memory updates. The division is intended to separate reusable motor primitives from the higher-level decisions that arrange them into a task.
Reported results
- When GPT-5.6 Sol learned skills on LIBERO-90 and GPT-6 Astra evaluated each frozen checkpoint, LIBERO-Pro Long success rose from 2.0% to 56.3%, despite no training on Pro Long.
- Independent Robosuite training produced mean success rates of 85.1% with Sol and 89.4% with Opus 5 across seven tasks.
- Frozen skills trained by Sol on LIBERO-90 reached 78.75% mean completion across four real-world manipulation tasks when executed with Astra.
- Removing the Verifier or Governor reduced final Pro Long success by 17.3 and 13.3 percentage points, respectively.
Why it matters
The paper’s main contribution is a change in emphasis. Instead of treating sim-to-real transfer as a requirement to relearn each end-to-end policy, it treats transfer as the movement of composable, inspectable capabilities through a common interface. The API provides an abstraction layer, the hierarchy separates motor execution from task composition, and the PVG loop uses information that is cheap to obtain in simulation but difficult to collect on a physical robot.
The results should still be read within their stated scope. They cover particular benchmarks, model combinations, and four real-world tasks; the quality of the API and the match between simulated and physical observations may strongly affect generalization. Broader robot embodiments, longer-horizon tasks, and more open-ended environments remain to be tested. Even so, Skill2Real presents a clear blueprint for building robot systems whose reusable unit is an executable skill memory rather than a single task-specific policy.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...