GraphForge Builds Verifiable Training Tasks for Working Agents
Introduction
A working agent must do more than respond to a prompt. It may need to inspect heterogeneous files, use several tools, coordinate actions across multiple steps, and produce a deliverable that another person can verify. Building training data for this behavior is difficult. Model-generated files often lack the messiness and diversity of real work, while tasks built on real files may not include reliable, task-specific verification.
GraphForge addresses this gap by treating the workspace and its evidence as the foundation of task synthesis. Rather than generating a task, files, and an evaluator independently, the framework attempts to connect all three through a structured evidence graph.
How the framework works
- Occupation-grounded seeds: The pipeline begins with seeds tied to occupational scenarios, providing controlled variation in the kinds of work represented.
- Real-file workspaces: For each seed, GraphForge assembles a workspace containing real files instead of relying only on synthetic documents.
- Evidence graphs: Relations among the files are represented in a graph. These relations supply the factual and procedural basis for the task.
- Traceable rubrics: Task requirements are derived from the graph, and each rubric criterion is linked to the files needed to verify it. This makes the evaluation path explicit rather than leaving correctness to a vague model judgment.
- Execution and repair: An initial rollout checks whether the task can actually be executed. A revision agent then repairs the instructions and rubrics against the original files before trajectories are collected.
Reported results
The paper reports that fine-tuning Qwen3.6-27B on 2,169 GraphForge trajectories produced a GDPVal score of 1,445.7 under OpenHands, a gain of 65.7. Under Claude Code, the resulting scores were 63.7 on Workspace-Bench-Lite and 24.0 on SpreadsheetBench II, improvements of 7.7 and 13.7 respectively.
The researchers also applied rejection fine-tuning to rollouts generated by the supervised fine-tuned model. Candidates were selected with the evidence-anchored rubrics, and the paper reports further gains on all three benchmarks. This result is important because it suggests that a rubric grounded in source files can serve not only as an evaluator, but also as a useful signal for choosing better training examples.
Why it matters
GraphForge connects four requirements that are often handled separately: realistic workspaces, executable instructions, verifiable outcomes, and scalable post-training data selection. Real files provide environmental complexity; the evidence graph constrains what the task can ask; execution checks expose broken tasks; and file-linked rubrics define how results should be inspected.
The available material does not establish how well the approach generalizes across more occupations, file formats, or longer workflows. Questions also remain around data rights and privacy, the cost of building evidence graphs, and the stability of automatic rubrics for open-ended deliverables. Still, the framework points to a useful principle for agent training: scaling task counts is not enough. The workspace, instruction, and verification evidence need to form a coherent loop.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...