Back to articles
Inference & Serving

When More Thinking Is Not Enough: FlyBy Teaches Small Models to Ask for Help

3 min read

Introduction

For reasoning models, more test-time computation often means generating a longer chain of thought. This is especially attractive for small reasoning models because they are inexpensive to serve. If additional thinking can recover difficult answers, deployment teams may avoid using a much larger model for every request. A new study from KAIST AI, however, asks a more fundamental question: does the model need more time, or does it need information that is not in its parameters?

The researchers intervene in intermediate reasoning states across two model families and several model scales. Their analysis suggests that self-refinement often concentrates probability on solutions that are already reachable from the current state. It does not necessarily create a new route to the answer. The finding does not invalidate extended reasoning, but it clarifies when extended reasoning is likely to help.

Two kinds of failure

The paper separates failures into two regimes:

  • Execution bottlenecks: The correct route is within the model’s existing capabilities, but the model makes an error while deriving, checking, or selecting a step. Reflection can sometimes recover the answer.
  • Knowledge bottlenecks: The model lacks a fact, concept, or piece of background information needed to continue. Generating more tokens may recycle uncertainty rather than make the missing route available. Relevant external information can change what the model is able to reach.

This distinction leads to a different inference policy. Internal computation is appropriate when the problem is execution. When the obstacle is knowledge, a targeted query to a stronger model may be more useful than another round of self-reflection.

How FlyBy works

FlyBy is a selective query-augmented reasoning framework built around this diagnosis. The small model first reasons on its own and then assesses what remains unresolved. If it detects a knowledge bottleneck, it can query a stronger external model whose parametric knowledge extends beyond its own. The returned information is then incorporated into the ongoing reasoning process. Queries are not issued indiscriminately for every problem.

Training has two main stages. Supervised fine-tuning bootstraps a multi-depth query action, allowing the model to learn different levels of external assistance. Cost-aware reinforcement learning then calibrates three decisions: whether to query, what to ask, and how much to spend. In this design, the stronger model is not merely a fixed answer engine. It acts as an on-demand extension of the smaller model’s capabilities.

Results and implications

Across 1,158 hard problems from six benchmarks, FlyBy-4B reaches 45.96% pass@8, compared with 41.64% for Qwen3-14B, while the reported serving cost is 2.7 times lower. Its pass@1 is 16.85%, above Qwen3-8B’s 15.31%. Scaling FlyBy to 8B raises pass@8 to 51.81%. These results suggest that model size, test-time compute, and external assistance do not have to be treated as mutually exclusive choices. A small model can handle routine reasoning and borrow stronger knowledge only when it encounters a genuine gap.

Selective querying is not free. The diagnosis itself introduces overhead, and the model must correctly interpret and integrate the returned information. FlyBy is therefore best understood as an inference orchestration strategy: identify the type of failure first, then allocate either more internal computation or external knowledge.

For cost-sensitive deployments, this is a useful direction. Future reasoning optimization may involve more than producing longer traces. Models may also need to learn when to continue, when to verify, when to ask, and when to stop. The advantage of a small model may ultimately come from managing its limited internal and external resources more intelligently.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles