The paper introduces Unbounded Positive Asymmetric Optimization (UP), a new importance sampling technique for reinforcement learning that overcomes the exploration‑stability dilemma by permitting unbounded positive policy updates while preserving training stability. Unlike standard clipping methods that restrict update budgets, UP enables sample‑efficient learning of complex reasoning in large language models.
Read original
huggingface/daily-papers