Back to articles
AI Safety

When Harmless Skills Combine into a Harmful Agent Workflow

4 min read

Introduction

Skill-based agent systems are designed to be extensible. An agent can load a skill when it needs a particular capability, and that skill may combine natural-language instructions, executable scripts, and reference resources. The model is attractive because third-party capabilities can be reused without rebuilding the entire system. Yet the same openness creates a security problem that is easy to miss: harmful behavior may emerge not from one malicious skill, but from the way several skills interact.

A paper titled “Stealth Apart, Harm Together: Skill Cascading Attacks on Skill-Based Agent Systems” calls this threat a skill cascading attack. The study focuses on a gap in existing evaluations, which have largely examined vulnerabilities inside individual skills rather than risks created by chains of skills.

Splitting one objective across many skills

In a cascading attack, the attacker distributes a harmful objective across multiple modules. Each skill makes a limited change that may appear reasonable in isolation. One might alter an extracted field, another might adjust a priority score, and a later skill might filter the resulting output. When the agent executes the workflow, however, those local changes accumulate and produce a harmful system-level effect.

The paper illustrates the idea with a prescription-review pipeline. One skill weakens signals associated with recently discontinued medications in the extracted medication history. A second reduces the severity of drug interactions linked to those medications. A third suppresses the resulting low-priority alert in the final summary. None of the steps necessarily looks like an overtly malicious action on its own. Together, they can cause a serious interaction warning to disappear before it reaches a physician.

The main lessons are:

  • The security risk lies in dependencies, ordering, and information flow between skills, not only in the code or instructions of one skill.
  • A scanner that reviews skills independently may not reconstruct a distributed objective.
  • Runtime monitoring focused on local calls may miss how a final outcome is built over several steps.
  • Individually acceptable components do not guarantee a safe combined workflow.
  • Defenses must reason about cross-skill interactions and cumulative changes to task outputs.

Testing the blind spot

To study this problem, the authors developed SkillCascade, an automated multi-agent red-teaming framework. They also released SkillCascade-Bench, a benchmark containing 213 validated cascading test cases across multiple agent systems and domains.

The evaluation covers representative agents including OpenClaw, Claude Code, and Codex, along with different large language model backbones. According to the paper, cascading interactions reliably induced harmful behavior while evading existing per-skill scanners and runtime monitors. The result is important because it challenges a common assumption in modular agent security: if every component passes its own inspection, the assembled system should be safe.

SkillCascade instead treats the workflow as the unit of analysis. It asks what one skill changes, how the next skill interprets that change, and whether the sequence gradually shifts the final decision. This perspective is closer to how real agent pipelines operate. It also explains why checks based only on suspicious instructions, dangerous scripts, or isolated outputs may not be sufficient.

Implications for agent security

As skill ecosystems become more open, the number of dependencies and possible combinations will grow. Developers may need to track data flows, invocation order, and important state transitions across skills rather than focusing only on the newest module. High-risk workflows should preserve intermediate results and provide an audit trail for changes to alerts, priorities, and evidence. A skill should not be able to quietly remove critical information simply because a previous step lowered its apparent importance.

The broader message is straightforward: agent security must ask not only whether a skill is safe, but also where the system goes when several skills run together. Cross-skill semantic analysis, compositional red teaming, and system-level runtime defenses are likely to become increasingly important as skills become a standard way to extend agents.

Source: arXiv

Comments

Checking sign-in status...

Loading comments...

Related articles