Paper page - Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex
…Among existing recipes, group-based policy gradient is prevalent, which samples a group of responses per prompt and updates the policy via group-relative advantage signals. This work reveals that these optimization…
