Back to articles
Large Language Models

How a Few Opening Tokens Can Unlock a Base Model’s Reasoning

3 min read

Introduction

Reasoning performance in large language models is often attributed to larger scale, chain-of-thought prompting, or reinforcement learning. A paper from researchers at MIT offers a more granular explanation: in some cases, whether a base model enters a reasoning-like trajectory may depend on only a few tokens at the beginning of its response.

The study calls these tokens token cues. They are not necessarily meaningful instructions. Instead, they reflect continuation patterns learned from the training corpus. If a particular opening frequently appears in documents containing detailed derivations, code explanations, or carefully structured answers, the model may associate that opening with the behavior that comes next.

Key findings

  • Fixed prefixes can change benchmark performance substantially. On MATH-500, forcing Olmo-3-7B to begin with “.\n\nOkay” raises pass@1 accuracy from 42% to 78%. For Qwen3-14B, the cue “Alright,” increases accuracy from 72% to 87%. These results suggest that the gap between a base model and an RL-trained model is not necessarily entirely due to newly acquired reasoning ability.
  • RL increases the likelihood of effective cues. The paper argues that reinforcement learning changes not only answer quality, but also how often the model selects openings associated with useful reasoning. Fixing those openings can recover a substantial portion of the improvement over the base model.
  • The effect can be traced to the data. Through causal interventions on training examples, the researchers turn an ordinary word such as “chicken” into an effective reasoning cue, or remove the effect of an existing cue. The result points to learned associations with document types and continuations rather than to the intrinsic meaning of the word.
  • Unusual instructions can be made effective. After a similar data edit, “Think duck duck goose” becomes as effective at eliciting reasoning as “Think step by step.” This suggests that familiar reasoning prompts work partly because of how they are represented in the training distribution.
  • Safety behavior is cue-sensitive as well. Different prefixes can produce distinct refusal and compliance patterns. Those patterns also correspond to different kinds of training documents, making token cues relevant to safety analysis as well as capability evaluation.

Why it matters

The study provides a useful way to think about the difference between base and reasoning-oriented models. A base model may already contain the capacity to generate extended problem-solving traces, but its default continuation may not enter the relevant trajectory. One role of reinforcement learning may be to make those trajectories easier to reach, rather than to create every underlying capability from scratch.

For model developers, this suggests that comparisons should test more than default generation. Prefix sweeps and response-start interventions could reveal whether an apparent capability gap is partly a difference in behavioral routing. For prompt engineers, effective openings may offer a low-cost performance boost, although such gains should not be confused with robust understanding and may vary across tasks.

The safety implications are equally important. If small, seemingly irrelevant tokens can alter refusal or compliance, then evaluation should examine these pathways explicitly. More broadly, the paper places stylistic and structural patterns in the training corpus at the center of reasoning research. Training may need to manage not only what models know, but also which textual cues activate reasoning, refusal, or execution behaviors.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles