Paper page - Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling
…We further identify that the conventional MTP training objectives ( cross-entropy or KL) are suboptimal in such settings, and therefore we propose a novel end-to-end TV loss that directly optimizes…