Paper page - STARE: Surprisal-Guided Token-Level Advantage Reweighting for Policy Entropy Stability
…Surprisal-Guided Token-Level Advantage Reweighting for Policy Entropy Stability Published on Jun 17 Submitted by haipengluo on Jun 18 Tencent Hunyuan Authors: Haipeng Luo , , , , , , Abstract GRPO algorithms face policy entropy collapse…