TTPO introduces Test-Time Policy Optimization to enable test-time training for large language models in mathematical reasoning. It addresses the fragility of using majority-vote pseudo-labels, which can corrupt model training when incorrect votes are used. The method aims to overcome the reliance on ground-truth labels required by traditional RL and On-Policy Self-Distillation.
Read original
huggingface/daily-papers