From Pareto Frontiers to Personal Preferences: A New Path for Test-Time Scaling
Test-time scaling (TTS) improves the reasoning performance of large language models by spending additional computation during inference. A system may sample multiple solutions, verify them, search over alternatives, or continue reasoning until a stopping condition is reached. These techniques can raise accuracy, but they also consume more time and inference resources. In practice, users rarely optimize only one of those dimensions.
A user may require a minimum accuracy while accepting a longer response time. Another may prioritize latency, even if that means using a less exhaustive search. A third may operate under a strict inference budget. This creates a gap between conventional TTS optimization and the requirements of real deployments. Existing work often advances one accuracy-efficiency frontier at a time, such as accuracy versus cost or accuracy versus latency. A point that looks attractive on one frontier may be unsuitable once several constraints are applied together.
The paper, “From Pareto to Preference,” formulates this setting as Personalized Test-Time Scaling. The objective is to discover executable controllers that maximize the rate at which a user’s requirements are jointly satisfied. In this context, a controller is a policy for allocating test-time computation: it can determine how much candidate generation, verification, or additional search should be performed and when the process should stop. The goal is therefore not simply to find the best average controller, but to identify a policy that fits a particular preference profile.
PersonTTS has several central components:
- Joint requirement satisfaction: Accuracy, latency, and inference cost are evaluated together instead of being optimized in isolation.
- Requirement-matched initialization: A new search can start from controllers discovered for similar user requirements rather than from scratch.
- Source-distilled procedural guidance: Earlier policy-discovery experience is converted into reusable guidance for the discovery agent.
- Target-profile evaluation: Although prior experience is reused, every candidate is still tested under the target user profile, preserving personalization rather than copying historical decisions blindly.
The authors evaluate the approach on AIME and HMMT, with experiments involving unseen user profiles and held-out problems. The reported results indicate that PersonTTS achieves higher joint requirement satisfaction than strong TTS baselines. When the number of candidate evaluations is held constant, reusing experience across users further improves policy quality and substantially reduces the time and cost spent by the discovery agent.
The broader implication is a shift in how TTS efficiency can be understood. Instead of searching for one universal operating point or merely presenting a Pareto curve, a serving system could discover and select policies according to the needs of different users, products, devices, or service tiers. This is particularly relevant when accuracy, response time, and inference spending are all first-class constraints.
The work also raises important evaluation questions. A TTS method should not necessarily be judged only by peak accuracy or a single cost metric; it may be more useful to measure how often it satisfies a complete requirement set. At the same time, the available summary does not provide numerical gains, and it remains to be seen how well the discovered controllers transfer across models, tasks, and changing workloads. Those questions require the full paper and broader validation.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...