Back to articles
Reinforcement Learning

Why Self-Training Can Hurt Reasoning Models: R-Quest Cleans Up the Questions

3 min read

Introduction

A self-evolving reasoning system can generate questions, solve them, and use the resulting data for another training round. This loop promises a way to improve reasoning without relying on a continuously expanding human-curated dataset. Yet self-training is not automatically self-correcting. Over time, a model may generate flawed tasks or keep producing slightly altered versions of problems it already knows.

The paper Questioning the Questions: Sustaining Self-Evolution in Reasoning Models examines why performance can deteriorate across repeated rounds and presents R-Quest, a framework designed to preserve useful diversity in the training stream.

Two failure modes in generated questions

The authors identify two recurring sources of data degradation.

  • Invalid questions become more frequent. A generated task may be underspecified, logically inconsistent, or otherwise impossible to answer properly. Filtering examples by answer consistency does not guarantee that the question itself is valid. According to the paper, this filtering strategy can even increase the share of invalid questions that survive into training data.
  • Mathematical repetition is hidden by wording changes. Lexical similarity methods can detect near-duplicate text, but they may miss two questions that use different wording while expressing the same underlying mathematical structure. Repeated variants gradually narrow the effective curriculum and can lead to question-diversity collapse.

Together, these findings shift attention from the solver alone to the question generator. A self-evolution system can only learn as broadly as the challenges it continues to receive.

How R-Quest changes the loop

R-Quest introduces two feedback signals: question validity and question novelty.

First, the solver is trained to recognize and reject invalid questions. Its judgments are then used in two places: to guide rewards for the questioner and to filter data used to train the solver. This makes invalid-task detection part of the evolution process rather than a one-time preprocessing step.

Second, R-Quest uses a frozen base model to compare sampled question pairs and provide novelty feedback. The fixed reference is intended to offer a more stable comparison point, helping the system identify cases where different surface forms represent the same mathematical problem.

Results and broader implications

The paper reports the strongest average performance for R-Quest across 12 benchmarks spanning mathematical reasoning, general-domain reasoning, and code generation, evaluated on two model families. Its longer-horizon behavior is especially notable: over ten self-evolution rounds, the method continues to improve, reaches its peak in the final round, and outperforms R-Zero by 17.32 points.

The central lesson is that answer agreement is not enough to validate self-generated training data. A robust loop must ask whether a task is well-formed and whether it adds a genuinely new challenge. Validity feedback limits the spread of erroneous supervision, while novelty feedback prevents the curriculum from collapsing into familiar templates.

R-Quest therefore frames self-evolution as a data-quality problem as much as a learning problem. Future systems may need even stronger measures of difficulty, information gain, and semantic duplication if they are expected to improve reliably over many rounds.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles