Back to articles
AI Safety

Same Text, Different Authority: How Control Tokens Amplify Prompt Injection

3 min read

Introduction

Prompt injection is often framed as a problem of malicious text entering a model’s context. This study highlights a lower-level factor: the tokenizer’s choice of representation. The same visible marker can be emitted as one reserved control token or split into several ordinary subword tokens. Both forms decode to the same text, but they do not present the model with the same learned vector. As a result, they may not have the same ability to make an injected instruction look like a genuine system or user turn.

Key findings

The researchers study forged chat-template markers, such as strings resembling <|im_start|>, inserted into untrusted content. They keep the decoded text unchanged and control for the additional tokens introduced by subword encoding, allowing the experiment to focus more directly on representation rather than text length.

  • On the InjecAgent benchmark, replacing forged reserved markers with ordinary subword sequences reduces attack success by 39 to 66 percentage points across three of four open-weight model families. The gap also transfers to multi-turn tasks in AgentDojo.
  • The difference is smaller on Qwen3-8B, at 8 points, because the model can still recognize a forged turn from its text through reasoning. When the reasoning block is suppressed, the gap expands to 50 points.
  • The embedding analysis finds that averaging the vectors of the marker’s subwords does not reproduce the reserved-token effect. On Llama-3.1, however, a nearby ordinary-token vector can restore the attack.
  • An adaptive attacker searching for non-reserved alternatives finds embedding neighbors in three of the four tested families.

Why it matters

The paper places much of the marker’s authority in a single learned vector at the marker position. This challenges defenses that treat control markers as ordinary strings or rely only on visible-text filtering. A string-level check may appear to remove the control signal while leaving other representations with similar model-level behavior.

For defenders, forcing suspicious markers through ordinary subword encoding appears to be a practical tokenizer-side mitigation. It is not a complete solution. A model may infer the intended turn boundary from the surrounding text, and an adaptive attacker may search the embedding space for a substitute that behaves like the reserved token. The Qwen3-8B result also shows that reasoning can partially compensate for the loss of a special representation.

The study further audits tokenizer configurations and reports that many still expose markers used by tool protocols. This matters because agents routinely process untrusted tool output through those same boundaries. Securing only direct user prompts therefore leaves a potentially important channel open. Tool results, protocol delimiters, multi-turn assembly, and privilege checks need to be considered together.

More broadly, the work expands the scope of prompt-injection research from “which strings are dangerous?” to “which internal representations grant authority?” Tokenizer design can reduce a substantial part of the risk, but robust protection still requires model-level training, strict separation of trusted and untrusted content, and careful control of tool permissions.

Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles

CCTest · Blog
When Do Model Internals Help LLM Safety? A Matched Comparison of DPO and Representation Engineering
AI Safety
cctest.ai
AI Safety

When Do Model Internals Help LLM Safety? A Matched Comparison of DPO and Representation Engineering

A matched study compares DPO, representation steering, internal probes, and text monitors across safety control and risk detection. Representation methods do not replace behavioral alignment, but they offer useful advantages in low-data and cost-sensitive settings.

Read more