HERMES Turns Repository Components into Collaborative Coding Agents
Introduction
Language models with terminal access can already edit files, run tests, and investigate bugs. Their reliability drops, however, when a task spans source code, configuration, dependencies, tests, and runtime behavior. The agent must repeatedly reconstruct a project’s state from scattered evidence. Long interaction histories then create context pressure and semantic drift, while large repositories make it difficult to identify the components that actually matter.
A paper featured on Hugging Face Daily Papers proposes HERMES, a framework that focuses on the organization of the agent’s working environment rather than only on the strength of its underlying model.
Dev-Primitives as an executable abstraction
The central idea is the Dev-Primitive, or Development Primitive. A repository artifact—such as a code component or test-related artifact—is paired with a resident language model. Instead of treating the artifact as a passive object inspected from outside, the paired model exposes an agent-native interface grounded in the artifact’s implementation and dependencies.
This design provides three main capabilities:
- Local reasoning: a primitive can explain behavior and constraints within its own scope, reducing the need to repeatedly load the entire repository.
- Natural-language communication: primitives can exchange state, assumptions, and proposed actions with other components.
- Localized revision: once evidence points to a component, the associated primitive can help formulate or apply a focused change.
HERMES instantiates these primitives at repository scale through dependency-aware dynamic activation. Rather than activating every component for every task, it selects a relevant subset based on the task and dependency structure. Its diagnosis mechanism then maps execution evidence, such as failures observed during testing or runtime, back to the components that should be revised. The intended loop is execution, diagnosis, and targeted modification.
Results and interpretation
The paper evaluates HERMES on four software engineering benchmarks against matched baseline harnesses. It reports an average improvement of 12.4%. With strong activation and diagnosis models, a configuration using Qwen3-8B as the Dev-Primitive models remains within 4.5% of a homogeneous GPT-5.6 Sol configuration across the four benchmarks. On Terminal-Bench 4.0, it also reduces inference cost by 26.2%.
These findings do not imply that a smaller model universally replaces a stronger one. They do indicate that task decomposition, context routing, and error attribution can materially affect the performance of coding agents. A capable model may still waste effort if the harness repeatedly exposes irrelevant repository state or fails to connect a runtime symptom to the right file or module.
Significance and open questions
HERMES reframes the repository from a static collection of artifacts into a cooperative environment made of executable, locally informed units. This could reduce the context burden of long-horizon work and make bug fixing more evidence-driven.
The approach also leaves important engineering questions open. The quality of results depends on the activation and diagnosis models, while primitive boundaries must be chosen carefully. Natural-language communication can introduce misunderstandings, and a local change still needs to preserve global consistency. Even so, the work points to a broader direction: improving coding agents may require designing better interfaces between models and software, not merely increasing model size.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...