Back to articles
Evaluation & Benchmarks

SWE-Game Tests Whether Coding Agents Can Build Playable Games

3 min read

Introduction

Getting a coding agent to launch a program is very different from asking it to build a game that follows a design, responds correctly to input, and remains genuinely playable. Games combine state management, collision handling, progression, feedback, and presentation. A missing rule or a small logic error can make an apparently complete project fail during play.

SWE-Game, featured by Hugging Face Daily Papers, is designed to measure this broader capability. Rather than checking only whether generated code runs, the benchmark asks whether an agent can understand a game specification, reproduce its mechanics, repair existing problems, and provide evidence that the result works.

Key findings

  • The benchmark covers a full development workflow. SWE-Game contains 247 tasks grounded in 41 executable Godot reference games. The games span 2D and 3D and cover 13 gameplay categories. Five task types address development from a brief, implementation from a game design document, skeleton completion, repair of 83 injected-fault cases, and porting from Godot to Unity.
  • Evaluation goes beyond video demonstrations. A shared instrumentation interface allows evaluators to drive independently implemented games and inspect runtime state through probes. Scores combine engine-state checks, certified replay of reference inputs, and agent-authored feature demonstrations. These components target mechanic correctness, demonstrated playability, and behavioral restoration or preservation after repair.
  • Presentation is scored separately. Game-specific vision-language rubrics evaluate visual presentation independently, separating the question of whether a mechanic works from whether the game looks polished.
  • Reliable construction remains difficult. Among six models, Opus5 achieves the highest overall score in all five task types. Even so, the best overall results for the three construction tasks stay below 60 out of 100, with Brief-to-Game reaching 50.38. Review of submissions identifies omitted requirements and gameplay logic errors as the dominant implementation problems.
  • Runtime evidence is especially valuable. On human-labeled behaviors from 100 agent-built games, executable checks achieve 92.59% balanced accuracy, compared with 78.41% for a video-based VLM judge. For visual quality, rubric-based scores achieve a Spearman correlation of 0.829 with human ratings across 200 gameplay clips.

Why it matters

SWE-Game is important not simply because it adds another leaderboard, but because it frames game development as a behavioral evaluation problem. Games provide a demanding test case: requirements are often expressed in natural language, correctness emerges through interaction, and failures may appear only after a particular sequence of inputs. Source-code inspection or a short video alone can miss these cases.

The findings suggest that current agents are limited not only by code-generation skill. They also have trouble tracking every requirement and coordinating related states and rules over time. A more dependable game-development agent will need to generate tests, explore edge cases, and use runtime evidence to confirm that a fix did not damage existing behavior.

For now, the practical role of these systems is closer to rapid prototyping, feature implementation, and targeted repair than unsupervised delivery of complete games. For benchmark designers, SWE-Game offers a useful template: combine reproducible interaction, engine-level checks, and visual assessment instead of relying on one signal alone.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles