ReToken introduces a learnable embedding to enhance vision-language models for visual retrieval by selecting query-relevant tokens from a pre-filled visual KV cache, addressing performance degradation with long visual contexts and GPU memory limitations. Trained on a small image-QA dataset, it achieves consistent