Back to articles
AI Agents

WEFT: Why Tool-Use Post-Training Needs a Whole System

3 min read

Introduction

Teaching a language model to use tools is not simply a matter of adding more APIs or executable environments. A useful training interaction also depends on the task being feasible, the agent harness exposing the right workflow, and the evaluator measuring completion correctly. If these components are misaligned, scaling one of them may produce little useful learning signal and can make system failures look like model failures.

WEFT, or Whole-system Evolution For Tool-use Post-training, addresses this issue by treating the full agentic interaction system as the object of scaling. The approach jointly considers environments, tasks, agent harnesses, and evaluators rather than optimizing the environment in isolation.

Core approach

  • Broader executable interactions. WEFT builds 8,172 executable MCPs, 64,755 tools, and 41,695 certified atomic tasks, which are composed into 11,884 longer tasks. The median task contains 20 atomic-task turns, and many tasks span multiple MCPs and domains.
  • More varied interaction formats. A grounded task can be presented as a complete brief for autonomous execution or revealed progressively by a simulated user. WEFT also routes tasks through native harnesses including ReAct, OpenClaw, and Hermes. With the task set fixed, mixing Agentic and SimUser data improves WEFT-35B-A3B by 3.18 percentage points on average. Adding the additional harnesses produces a further 9.71-point mean gain across the reported benchmarks.
  • Execution as a debugging signal. A failed rollout is not automatically a policy failure. WEFT examines tool traces, checkpoint results, database changes, and workspace state to identify whether the responsible component is the model policy, environment, task, or verifier. The suspected component is revised, then fresh rollouts test the change. Across three self-evolution rounds, selected teacher-trajectory tool-call errors fall from 1.76% to 0.96%, while the reported benchmark gains are 5.25, 4.33, and 3.65 percentage points.
  • More stable long-horizon training. Prefix-preserving sampling keeps verified progress and retries from the failed atomic task instead of discarding the entire trajectory. Atomic-turn credit assignment compares candidate segments from the same history and state, linking the signal to the current atomic task. Before reinforcement learning, a language model is used to filter tasks whose rubric judgments align with executable verifier scores; the RL process itself still uses executable rewards.
  • Recoverable concurrent state. MegaMCP hosts reusable services outside agent sandboxes while keeping each rollout’s database and workspace private and recoverable. Snapshots allow retries and branching without replaying completed interactions. In the reported 1,000-task experiment, sandbox upload volume drops from 773.5 to 34.8 MiB, a 95.5% reduction.

Why it matters

WEFT presents tool-use post-training as a systems problem. Task quality, interaction design, harness behavior, and verification all influence whether a rollout produces a meaningful learning signal. This is particularly important for workflows that cross tools, domains, and many dependent steps.

The method also reframes execution traces. They are not only examples for policy learning; they are evidence for improving the training system itself. Preserving successful prefixes and isolating recoverable state can reduce the cost of retries, while atomic credit assignment limits the spread of noisy rewards across long trajectories. The reported results are tied to the paper’s models, tasks, and benchmarks, so broader transfer to other models and production settings remains an open question.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles