RealCompanion Tests Whether AI Can Understand People Over Time
Introduction
An AI companion that speaks with someone for months should do more than archive old messages. It should gradually form a usable understanding of the person: what they have said, who they are, and whether something from the past is relevant to the message in front of it. RealCompanion is designed to test that broader capability in a setting closer to an ongoing relationship than a conventional memory question-answering benchmark.
What the benchmark contains
The release covers 10 human–AI relationships, with as many as 120 days of interaction and 27,218 messages in total. Alongside the raw conversations, the authors provide four derived resources: a profile, a persona, chat ground truth, and a question set. Each item cites the messages on which it depends. Chat labels also include a reasoning trace that was checked step by step against the conversation, making it possible to inspect not only an answer but the evidence behind it.
This structure matters because many existing evaluations decide in advance what a model should remember and then ask direct questions about it. RealCompanion instead keeps the full conversational setting in view. The challenge is to determine whether a message calls for historical evidence, locate the right evidence, and integrate it appropriately rather than merely producing a fact that appeared somewhere in the transcript.
Three findings
- Memory is rarely required. When all probes are pooled, a recency window can find the required message for 95.9% of them. Among probes that genuinely need memory, however, the same figure is only 2.2%. At the natural rate of different probe types, 96% of the gain from supplying recorded evidence comes from messages that did not need memory at all. Aggregate scores can therefore give a misleading picture of memory performance.
- Memory triggering is harder than retrieval. None of the tested detectors could reliably tell when memory was needed on real messages. Questions authored over the same histories leaked the cue more effectively, suggesting that benchmark construction itself can make the task easier. Simply labeling messages as memories increased their use by 10 to 14 percentage points, indicating that systems may often fail to recognize relevance rather than fail to access information.
- Similar quality can hide very different costs. Three agent systems achieved the same F1 score when reconstructing the persona, while their costs differed by a factor of 31. Accuracy alone is therefore insufficient for comparing long-term companion architectures.
Why it matters
RealCompanion points to a gap in how conversational memory is usually evaluated. Loading more history into a context window, or asking more direct recall questions, does not necessarily measure whether a system understands a person. A stronger evaluation should separate at least three abilities: understanding the current message, deciding whether the past is relevant, and selecting and using the right evidence without introducing unrelated history.
The result also has practical implications for product design. A useful memory system needs more than storage and search. It needs a trigger policy, evidence grounding, and a cost-aware way to decide how much context to process. Overuse of old messages can make an assistant feel intrusive, while underuse prevents a relationship from becoming coherent. The central challenge is not simply remembering more, but remembering at the right moment.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...