The authors introduce RealCompanion, a benchmark that evaluates how well an AI companion can understand a human partner by reasoning over extended, real‑world dialogues. They collected ten longitudinal relationships between users and an AI companion, comprising 27,218 exchanged messages spanning up to 120 days. For each relationship they released the raw conversation alongside four derived artefacts—a user profile, a persona description, a chat‑level ground truth, and a question set—each explicitly linked to the source messages that justify it. Every chat label is accompanied by a stage‑by‑stage reasoning trace that is validated against the dialogue. Analysis of the benchmark yields three primary findings. First, historical context is seldom required: a simple recency window retrieves the needed message for 95.9 % of all probes, yet only 2.2 % of probes that actually depend on memory are captured this way, and 96 % of the performance gain from providing full history stems from messages that impose no memory demand. Second, no existing detector can reliably identify when memory is necessary on authentic messages; moreover, questions generated from the same histories inadvertently reveal the answer, and manually tagging messages as “memory” increases their utilization by 10–14 percentage points. Third, three distinct agent architectures achieve comparable persona‑reconstruction F1 scores while differing in computational cost by a factor of 31, highlighting a substantial trade‑off between effectiveness and resource consumption.
Read original
huggingface/daily-papers