MemFold introduces a compact soft memory mechanism for long-context personalization that optimizes a fixed-budget latent interface through on-policy reinforcement learning rather than reconstruction or imitation objectives. The method compresses query-conditioned textual memory into K continuous vectors serving as the reader's memory interface, then trains the reader on its own rollouts using two complementary signals: group-relative rewards for task outcomes and confidence-gated on-policy distillation. In the distillation component, a frozen textual-memory teacher re-scores the student's sampled tokens under the full textual memory, providing supervision that remains on the student's current distribution without requiring autoregressive decoding from the teacher, which is discarded at inference. Evaluated across three Qwen backbone models, MemFold achieves the highest reported accuracy on PersonaMem-32K and PersonaMem-128K benchmarks, with performance gaps widening at the longer 128K context length. The approach transfers zero-shot to PrefEval and LongMemEval without target-domain fine-tuning. Ablation studies attribute the majority of task improvement to the reward signal, with a smaller incremental gain from the teacher distillation term, while memory intervention experiments confirm the reader relies on instance-specific content encoded in the soft memory vectors.

Read original