The authors propose a method to train large language model (LLM) judges using natural language feedback through position‑selective self‑distillation. Unlike outcome‑supervised reinforcement learning approaches such as GRPO, which assign a single scalar reward to every token based solely on the correctness of the final verdict, their technique treats the same model, conditioned on the feedback, as a teacher that supplies dense, token‑level supervision. They compute the per‑position entropy shift between the teacher’s and student’s output distributions to distinguish two regimes: context sharpening, where the teacher concentrates probability on a specific feedback‑aligned criterion expression, and context spreading, where probability is dispersed across multiple feedback‑aligned alternatives. Interpreting sharpening as encouraging memorization and spreading as fostering semantic understanding, they devise a position‑masking strategy that retains the lower tail of the entropy‑shift distribution (i.e., masks higher‑entropy‑shift positions). Experiments demonstrate that this selective masking improves out‑of‑distribution generalization compared with naïve self‑distillation. The resulting self‑distilled judges exceed outcome‑supervised RL‑trained judges by 2–9 percentage points on subjective evaluation subcategories while maintaining comparable performance on objective tasks.
Read original
huggingface/daily-papers