SLCA-GRPO Gives Tool-Calling Agents More Precise RL Credit
Introduction
Tool-calling agents do not generate one homogeneous kind of output. Their trajectories may alternate between structured function calls and natural-language explanations or summaries. This creates a training challenge for policy optimization methods such as GRPO. In the conventional setup, a scalar advantage computed for an entire rollout group is broadcast to every token in the trajectory. A reward signal associated with the final summary can therefore influence tool-selection tokens, while execution-related variation can also affect the training of user-facing language.
The paper introduces SLCA-GRPO to address this problem at the level of credit assignment. It also presents the Schema-Guided LLM Simulator, or SGLS, as infrastructure for exploration and training without repeatedly calling expensive real-world APIs.
What the method changes
- Segment-level advantages: Segment-Locked Credit Assignment (SLCA) estimates advantages for structural segments within the same rollout group instead of using one undifferentiated trajectory score.
- Reward routing: With Hierarchical Rewards (HierR), execution advantages are assigned to tool-call tokens, while preference-related advantages are assigned to summary tokens.
- No extra intermediate rollouts: The approach does not require collecting new rollouts from intermediate states, preserving the basic group-based sampling process.
- Schema-guided simulation: SGLS supplies a structured simulated tool environment, making large-scale exploration more practical and reducing dependence on costly external services.
Reported results
On a 7B backbone, the paper reports faster convergence and better results than standard GRPO, ToolPO, and RLTR under the same training budgets. The reported gains are 2.53 percentage points on in-domain evaluation, 1.36 points on the Berkeley Function-Calling Leaderboard, and 9.15 points on τ²-Bench. Additional results across three Qwen backbones and both in-distribution and out-of-distribution tool-calling benchmarks indicate higher task success and fewer tool turns.
Why it matters
The broader lesson is that tool-calling policies should not treat every token in a mixed-format trajectory as if it had the same learning role. Choosing a tool and explaining the result are related, but they are evaluated through different signals. Broadcasting one advantage across both segments can inject noise into the policy update; locking each signal to its corresponding segment offers a more targeted alternative.
The approach also depends on the quality of its reward hierarchy and simulator. SGLS may not capture every failure mode found in real APIs, and the stability of segment-specific rewards in longer, more complex tasks remains an open question. Even so, SLCA-GRPO provides a focused way to improve optimization efficiency while reducing unnecessary tool use.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...