Back to articles
AI Agents

AdvSim2Real Trains Web Agents Against Evolving Prompt Injection

3 min read

Introduction

Prompt injection is especially difficult for web agents because the content they must trust and the content they must distrust can appear on the same page. A shopping site, document, or form contains the values and controls needed to complete a user request, yet a third party may insert text such as “ignore the original task” or “perform a different action.” Treating all page content as irrelevant would make the agent ineffective; following it indiscriminately can make the agent unsafe.

AdvSim2Real addresses this tension through co-evolution rather than a fixed collection of attack examples. The authors place a task generator, an injection adversary, and the web agent inside a frozen web world model. Training takes place in this simulated environment, with the goal of transferring the resulting capability and robustness to an actual browser.

Key points

  • An adaptive curriculum keeps producing useful tasks. The curriculum is rewarded for tasks the agent completes about half the time. Easy tasks quickly stop providing a learning signal, while extremely difficult tasks offer little actionable feedback. Focusing on the boundary between success and failure keeps training challenging without making it uninformative.
  • The adversary optimizes for outcome changes. It is not enough for an attack to look suspicious. The injection must turn a run that would have been judged successful into a failure. This ties adversarial learning directly to the user’s task outcome.
  • The three components evolve together. The agent trains against newly generated attacks as well as previous ones, while the task distribution changes as the agent becomes stronger. This reduces the risk that a fixed benchmark becomes easy and stops teaching.
  • Capability and robustness both improve. Across 150 web tasks, clean completion for the 4B agent rose from 74.89% to 81.33%, and attacked completion increased from 48.07% to 57.48%. Against an unseen Kimi-K3 attacker, performance rose from 23.00% to 30.72%, a 33.6% relative improvement. Strict success in a real-browser evaluation increased from 25.56% to 44.44%.

Why it matters

The main contribution is a change in the training signal. Instead of collecting a static list of malicious instructions, the system continually searches for tasks worth learning and attacks that actually defeat the current policy. That setup is closer to the adaptive nature of real-world prompt injection, where attackers can change wording and placement after observing a system’s behavior.

The results should still be read within the limits of the reported evaluation. A simulated web world cannot represent every site layout, tool interaction, or long-horizon failure mode. Transfer to a real browser is encouraging, but it does not by itself establish deployment-level security. Broader tests will be needed across websites, tools, task types, and attacks that exploit more than plain text.

Even with those caveats, AdvSim2Real suggests a practical direction for agent security: use a world model to run inexpensive, repeatable adversarial training, and keep the challenge moving as the agent improves. Robustness then becomes a continuing curriculum rather than a one-time fine-tuning stage.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles