Paper page - Current World Models Lack a Persistent State Core
… Implement Phases 0–2 instrumentation, α-ramp, η-attention . Go/no-go threshold: α can reach ≥0.5 with held-out perplexity within ~2% of baseline AND η-norm drift iii stays <1%. …
… Implement Phases 0–2 instrumentation, α-ramp, η-attention . Go/no-go threshold: α can reach ≥0.5 with held-out perplexity within ~2% of baseline AND η-norm drift iii stays <1%. …
… We introduce Self-Evaluation Elicitation SEE , a method that surfaces this latent ability through a short cycle comprising a calibration-coupled reinforcement learning phase that improves the answer and predicts the judge, followed by a masked distillation phase that sharpens the prediction while l… …
… Nice work! the arc-grounded context architecture is neat, but i'm curious how it handles non-linear character arcs where phases loop back or diverge after long lulls. what happens when arcs loop back or reset after big plot twists, and does that hurt phase fidelity when the text lets the character … …
… EFT thus serves as a "practice phase" for general-purpose discovery agents that do not solve new problems from scratch. …
Papers arxiv:2606.19162 The Reward Was in Your Data All Along: Correcting Flow Matching with Discriminator-Guided RL Published on Jun 17 Submitted by Nicolas Beltran-Velez on Jun 18 AI at Meta Authors: Nicolas Beltran-Velez , Felix Friedrich , , , , , Abstract Discriminator-Guided Reinforcement Lea… …
… Across all six directions of {Qwen3-4B, 8B, 14B} and six in-domain and out-of-domain benchmarks, our method outperforms prior heterogeneous baselines, matches or exceeds text communication in context-aware settings at roughly 2 to 3 times lower compute, and remains effective in context-unaware tran… …
… OpenAI's evaluation was probably running with the guardrails disabled. unless i missed it somewhere, this is a giant block of text and none of it includes how these companies should be working together to make sure THIS DOES NOT HAPPEN. it's basically a full circle moment when the discussion was al… …
… The core idea of BR-MoE is a bi-level routing mechanism: a router selection stage that dynamically activates relevant task-specific routers , followed by an expert routing phase that dynamically activates and aggregates experts, aiming to inject discriminative and comprehensive representations into… …
… The following papers were recommended by the Semantic Scholar API Diagnosing Training Inference Mismatch in LLM Reinforcement Learning 2026 AIS: Adaptive Importance Sampling for Quantized RL 2026 Reformulate LLM Reinforcement Learning for Efficient Training under Black-box Discrepancy 2026 Rethinki… …
… The following papers were recommended by the Semantic Scholar API PEEK: Predictive Queue-Informed KV Cache Management for LLM Serving 2026 Lodestar: An Online-Learning LLM Inference Router 2026 Recency/Frequency Adaptive KV Caching for Large Language Model Serving 2026 ModeSwitch-LLM: A Lightweight… …