Paper page - Rethinking the Divergence Regularization in LLM RL
…advantage-weighted quadratic regularizer on policy shift . DRPO preserves the same trust-region geometry as DPPO while inducing bounded, continuous gradient weights that attenuate diverging updates and provide corrective signals beyond the…