Back to articles
Robotics & Physical AI

RoboFoundry Turns Robot Experience into an Evolving System

3 min read

Introduction

A robot completing one task does not necessarily mean that it has learned how to handle the same class of tasks in the future. Many embodied-agent systems optimize memory, context management, skill libraries, or action interfaces as separate components. The foundation model makes decisions, while surrounding modules provide information and execute actions. Yet if failures are not converted into reliable system changes, the next attempt may effectively start from scratch.

RoboFoundry proposes a different view. Its central idea is Self-Evolving System-as-Policy: the policy of an embodied agent is not limited to model parameters or a single prompt. It also includes the way the system manages context, stores memory, organizes skills, and recovers from failures.

Key ideas

  • Diagnosing gaps from execution traces. RoboFoundry examines task execution to identify weaknesses in decision-making and memory management. The resulting experience is turned into task-specific system updates rather than being stored as unstructured logs alone.
  • Two evolving surfaces. A context system manages active internal information and persistent file-system memory. A hierarchical skill system organizes atomic skills, reusable compositions, and recovery procedures conditioned on failure modes.
  • Validation before generalization. Newly created capabilities are first tested on the relevant task. Improvements that recur and prove reliable can then be promoted to the general system, reducing the chance that a single bad experience becomes permanent behavior.
  • Separating decisions from embodiment. A shared semantic interface distinguishes embodiment-invariant decisions from robot-specific execution. This design is intended to let evolved capabilities transfer across heterogeneous platforms.

Reported results

On EmbodiedBench, the paper reports state-of-the-art performance, including a 27.8% improvement over GPT-5.5. Qwen3.7-Plus reaches 70.3%, approaching GPT-5.5 at 72.7% under the reported setting. This suggests that the framework’s benefits are not restricted to one foundation model.

For long-horizon memory, RoboFoundry outperforms every baseline on RoboMemArena by at least 39.0%, including methods assisted by external foundation models. On LIBERO-PRO, it reports improvements of 243.8% to 679.7% over Cap-Agent0 across the listed perturbation types. The paper also describes zero-shot transfer and online evolution in real-world robot deployments.

Why it matters

The broader contribution is a shift in how self-improvement is framed. Instead of treating every failure mainly as data for later model training, RoboFoundry allows it to modify the agent’s memory structure, skill hierarchy, and future recovery procedures. This is particularly relevant for long-horizon tasks, where failures can arise from planning, forgotten information, or poor coordination between abstract decisions and physical actions.

The available material does not provide the full update algorithm, the cost of validation, the rate of harmful updates, or detailed comparisons across hardware platforms. Whether this approach remains stable in more open-ended and safety-critical environments therefore requires further evidence. Still, RoboFoundry presents a compelling direction: embodied intelligence may advance not only through larger foundation models, but also through systems that can filter, verify, and retain useful experience.

Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles

CCTest · Blog
Teaching Robot Hands to Operate Like Humans: Morphometric Imitation Bridges Retargeting and Zero-Shot Deployment
Robotics & Physical AI
cctest.ai

Teaching Robot Hands to Operate Like Humans: Morphometric Imitation Bridges Retargeting and Zero-Shot Deployment

Researchers from UC Berkeley and collaborators introduce Morphometric Imitation, a three-stage pipeline that turns reconstructed human hand-object interactions into executable robot demonstrations and visuomotor policies. Across three robot hands and ten interaction tasks, the method reached a 89.3% zero-shot success rate in real-world trials.

Read more