Search

Showing top 11 results for "Disc phase-out"

huggingface.co › papers › 2606.29526

Paper page - The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning

… The following papers were recommended by the Semantic Scholar API Diagnosing Training Inference Mismatch in LLM Reinforcement Learning 2026 AIS: Adaptive Importance Sampling for Quantized RL 2026 Reformulate LLM Reinforcement Learning for Efficient Training under Black-box Discrepancy 2026 Rethinki… …

Jul 6, 2026