The authors observe that in online reinforcement learning, prompts vary widely in informativeness—some are already saturated while others are too difficult—yet all receive equal rollout budget. They introduce an exploration‑guided prompt scaffolding framework that dynamically adapts the training prompt distribution during RL post‑training of multimodal large language models. This approach aims to allocate more compute to informative prompts and reduce waste on unhelpful ones.
Read original
huggingface/daily-papers