The article introduces a synthetic training pipeline to enhance large language models' (LLMs) context-learning capabilities, addressing the challenge of memorization when using public documents already present in pretraining data. The method involves four steps: (i) rewriting source documents to reduce memorization risk, (ii) generating questions requiring reasoning over the document, (iii) answering these questions using the document as context, and (iv) filtering samples to ensure genuine dependency on the document. This pipeline produced approximately 10,000 training samples from 3,500 documents without human annotation. When applied to a student model (Qwen3.6-35B-A3B), supervised fine-tuning (SFT) improved its CL-bench score from 13.7% to 22.8%, with a subsequent rubric-reward reinforcement learning (RL) stage reaching 24.6%, nearing the performance of a frontier model (Qwen3.8-2.4T at 23.9%). The approach also demonstrated broad transfer to long-context understanding, instruction following, and reasoning, though code generation and knowledge remained largely unaffected. This scalable, reproducible method aims to advance context-grounded reasoning in LLMs while mitigating memorization risks.
Read original
huggingface/daily-papers