Paper page - The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning
…The following papers were recommended by the Semantic Scholar API Diagnosing Training Inference Mismatch in LLM Reinforcement Learning (2026) AIS: Adaptive Importance Sampling for Quantized RL (2026) Reformulate LLM Reinforcement Learning for…