LLMs Get Lost in Evolving User Intent

LLMs Get Lost in Evolving User Intent

Jihoon Tack, Philippe Laban, Jennifer Neville 2026-07-21

Researchers from Hugging Face and collaborators examine how large language models (LLMs) handle evolving user intent in multi-turn interactions, finding that current evaluation and training methods—largely based on singl…

→ View original source
Predictive Divergence Masks for LLM RL

Predictive Divergence Masks for LLM RL

Xiangxin Zhou, Jiarui Yao, Penghui Qi, Bowen Ping, Jiaqi Tang 2026-07-11

Reinforcement learning for large language models (LLMs) typically relies on trust-region masks to stabilize off-policy updates. The dominant PPO-style approach uses the sampled-token importance ratio

→ View original source
Loading more articles...