UNREAL introduces a model‑internal evidence selection mechanism that works for both corpus‑scale retrieval and long‑context inference. It extracts retrieval queries from the frozen LLM’s internal representations by encoding document chunks, requiring fewer than 500 K additional trainable parameters while leaving the backbone model untouched. Evaluated on a 3‑billion‑token, 21‑million‑chunk Wikipedia index, all four dense and hybrid UNREAL backbones surpass existing retriever‑reranker pipelines. The strongest configuration lifts HotpotQA recall from 49.1 % to 73.2 %, 2WikiMultiHopQA recall from 31.7 % to 60.1 %, and MuSiQue recall from 8.8 % to 14.4 %. When applied to long‑context settings, the same selector prunes distractors before generation, boosting NoLiMa accuracy from 1.0 % to 24.83 % at a 128 K‑token context and raising LV‑Eval F1 from 49.97 % to 54.66 % at 256 K tokens. Moreover, UNREAL cuts FLOPs and time‑to‑first‑token relative to full‑context processing once the input exceeds roughly 32 K tokens, with savings increasing as context length grows. These results demonstrate that a unified, model‑native selection approach can serve as a common foundation for both large‑scale retrieval and evidence‑sparse long‑context reasoning.
Read original
huggingface/daily-papers