Do Personalized LLMs Invent User Profiles? A New Benchmark Says Yes
Personalized LLMs are increasingly used with persistent memory, but there has been little systematic evidence about whether their user models are faithful. This paper asks a practical question: when a model tries to personalize responses, does it stay grounded in the available evidence, or does it fill in missing details by guessing?
To study this, the authors introduce MirageBench, a benchmark built around 150 personas spanning stereotypical, counter-stereotypical, and neutral profiles. It includes 6 personalization tasks that range across an imagination gradient, from conservative inference to more speculative settings. An independent judge was used to label 143,616 claims, and the setup was validated against blind human annotation.
Key takeaways
- Over-inference is widespread. Every one of the 12 evaluated models over-inferred on 35%–49% of its claims, with a cross-model mean of 41.6%.
- The problem depends on the task. OI rates varied from 27% to 59% across tasks, showing that personalization context matters a lot.
- Errors accumulate over turns. In a multi-turn pilot, inferred attributes tended to grow almost linearly, with little correction once a guess had been made.
- Self-monitoring can be inverted. Across models, self-reported OI was negatively rank-correlated with judge-measured OI. In other words, the models that claimed to over-infer the least were often the ones found to fabricate the most.
- Within-model self-audits still carry some signal. The models could rank their own claims moderately well, but that does not make self-report a reliable basis for comparing different systems.
Why this matters
The findings challenge a comforting assumption in product design: that a model’s own confidence or self-checking can stand in for external trust evaluation. For personalized assistants, the risk is not only factual hallucination, but also the fabrication of stable user traits that may shape future responses, recommendations, and memory updates.
The paper’s broader message is simple: personalization needs verification. If a model is going to store or infer user attributes, those inferences should be auditable and externally checked rather than trusted because the model says it was careful.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...