Back to articles
Evaluation & Benchmarks

A2Z GameSpec-Bench Tests Whether Coding Agents Follow Game Designs

3 min read

Introduction

A coding agent can now produce a game that launches and appears plausible. That, however, is a weaker achievement than faithfully implementing a complete game design. Long-form Game Design Documents (GDDs) connect rules, visual presentation, game state, and player actions. A local implementation may look correct while breaking a relationship elsewhere in the design. A2Z GameSpec-Bench, introduced by a KRAFTON research team, is designed to measure this gap.

What the benchmark evaluates

  • Long-form specifications: The benchmark contains 100 long-form GDDs. They are intended to represent a more demanding development setting than short prompts that describe only a few isolated features.
  • Dependency-aware contracts: Each GDD is converted into a contract containing requirements, constraints, and prerequisite relationships. This representation makes explicit which conditions must hold before another rule or interaction can work.
  • Evidence from code and play: Evaluation combines source-code inspection with agent-generated test policies. These policies drive scenario-based replay and adaptive playtesting, allowing the benchmark to examine implementation, runtime or rendered behavior, and player interaction together.
  • A stable basis for comparison: The contract remains fixed across agents and revision rounds. Judgments and evidence are linked to the same requirements, making it easier to compare systems and identify recurring violations instead of changing the target during evaluation.

Findings

The reported results show that current coding agents have difficulty satisfying interdependent requirements simultaneously in code and in actual play. Compilation and execution therefore provide an incomplete picture. A generated game may start successfully and still fail when a particular game state, sequence of actions, or interaction exposes a broken dependency. Static inspection alone can miss these failures because some violations only become visible during replay or playtesting.

The study also compares revision strategies. After two rounds, requirement-specific feedback improves GDD Fidelity by 10.9% relative to self-revision. The result suggests that agents benefit more from evidence tied to a concrete requirement than from a general instruction to inspect and improve their own output.

Why it matters

A2Z GameSpec-Bench shifts the evaluation question from “Can an agent generate executable code?” to “Can it preserve a complex specification through implementation and use?” Games provide a particularly concentrated test because logic, presentation, state, and interaction are tightly coupled. The same principle applies to other generated applications: when requirements depend on one another, a successful build or a single convincing demo is not enough.

For developers, dependency-aware contracts offer a structured way to connect failures with the original design. For benchmark researchers, fixed contracts and requirement-level evidence create a more consistent basis for comparing agents and revision methods. The project’s code and datasets are available publicly, providing a foundation for further work on specification-following in end-to-end software generation.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles