Back to articles
AI Safety

SkillDRE Connects Pre-Execution Scanning and Runtime Feedback

3 min read

Introduction

As agents increasingly rely on reusable skill packages, their security boundary extends beyond the model itself. A skill can bundle natural-language instructions, executable code, and task-specific resources. That combination can improve task completion, but it can also provide an adaptable container for concealed malicious behavior. The paper SkillDRE examines how such skills may evolve when attackers receive feedback from more than one defense layer.

Key ideas

  • A cross-stage feedback loop. Many evaluations treat pre-execution scanning and runtime protection as separate checks. Yet a skill that passes a scanner may fail when runtime defenses intervene, while a revision that improves execution may create new scanner findings.
  • Scanner and runtime signals are used together. SkillDRE first constructs a task-conditioned malicious objective and a verifiable judge rule from a benign task and its associated skills. The objective and judge remain fixed while the implementation is repeatedly revised according to scanner results and execution outcomes under runtime defense.
  • Runtime revisions return to scanning. Runtime feedback does not directly produce a final submission. Each revision is sent back through the pre-execution stage for rescanning and further optimization before another execution attempt, creating a closed loop across the two stages.
  • Benign functionality is part of the constraint. The framework also seeks to preserve legitimate task capability, making the evaluation closer to a disguised or capability-preserving threat rather than a purely disruptive payload.

Results and implications

On SkillsBench, evaluated across four victim models, SkillDRE reports a 45.28% average attack success rate, 40.3% higher than the strongest baseline. The final submitted skills received no SkillScan findings and largely retained benign-task performance. These figures do not imply that every evolved skill reliably bypasses defenses. They do show that an attacker able to observe multiple defense stages may exploit gaps between them, so a clean result at one stage is not necessarily a reliable estimate of end-to-end risk.

For defenders, the paper points to the need for evaluations that cover the full path from submission-time inspection to actual execution. Security systems should also account for the possibility that a skill changes in response to feedback, rather than treating each candidate as an isolated static sample. Useful safeguards include version tracking, sandboxing, behavioral auditing, and cross-iteration analysis.

The broader contribution of SkillDRE is methodological. Once reusable skills can be revised through feedback, defenses are no longer facing a fixed artifact but an adaptive object. Future work will need controlled red-team protocols, reproducible benchmarks, and audit mechanisms that test this adaptation without exposing unnecessary defensive details.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles

CCTest · Blog
When Do Model Internals Help LLM Safety? A Matched Comparison of DPO and Representation Engineering
AI Safety
cctest.ai
AI Safety

When Do Model Internals Help LLM Safety? A Matched Comparison of DPO and Representation Engineering

A matched study compares DPO, representation steering, internal probes, and text monitors across safety control and risk detection. Representation methods do not replace behavioral alignment, but they offer useful advantages in low-data and cost-sensitive settings.

Read more