Policy optimization for LLMs faces a stability-exploration trade-off mediated by Policy-KL regularization, which constrains response behavior and consumes the exploration budget when enforced, yet leaves optimization without explicit drift control when removed. The authors propose environmental regularization, shifting regularization to the input side to break this dilemma. As training progresses, the distribution of inputs is regularized to maintain stability while preserving exploration capacity.
→ View original source
huggingface/daily-papers