Paper page - Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle for Language-Model Post-Training
…learning and dense supervision, using sparse rewards for teacher model discovery and dense rewards for student model compression. AI-generated summary In settings where labeled verifiable training data is the binding constraint…