Paper page - GORGO: Online Tuning for Cross-Region Network-Aware LLM Serving
…During held-out evaluation windows, we fix the parameter values learned from tuning and improve p95 TTFT by 6.9-15.5% and p95 end-to-end (E2E) latency by 14.3…
…During held-out evaluation windows, we fix the parameter values learned from tuning and improve p95 TTFT by 6.9-15.5% and p95 end-to-end (E2E) latency by 14.3…
…The resulting system achieves real-time 1280 x 704 resolution editing at 24 end-to-end FPS on a single RTX 5090 GPU, with the DiT core running at 58 FPS. Experimental…
…This approach integrates the statistical properties of CPCA into an end-to-end trainable framework, enforcing the discovery of a shared subspace across diverse domains while preserving interpretability. Experiments on four standard…
…Generated by thinkingmachines/Inkling-Small Pixel-space diffusion models aim to learn an end-to-end generator directly over raw pixels. This is challenging because a single model must capture both global…
Papers arxiv:2606.12575 High-Fidelity Two-Step Image Generation via Teacher-Aligned End-to-End Distillation Published on Jun 10 Submitted by Dongyang Liu (Chris Liu) on Jun 12 Tongyi-MAI…
…agents mainly tune local policy variants rather than explore new policy families, even when given strategy guides and paper references. Scaffolds requiring each iteration to cite, instantiate, and adapt a prior method…
…Kangsheng Duan , Ziyang Xu , , , , Abstract A lightweight image inpainting framework achieves high-fidelity results with significantly reduced parameters and inference time through novel local-global interaction blocks and adaptive distillation strategies. Generated…
…To this end, we propose SAAS, a novel RL framework designed to cultivate dynamic self-awareness that precisely regulates search behavior without compromising accuracy. SAAS introduces three key components: (i) a search…
…We further improve this single-rollout strategy with practical value-model training designs. To improve optimization stability, we introduce a strict double-side token-level clipping strategy. SAO is able to train…
…End-to-End MultiModal Customization for Joint Audio-Video Generation (2026) Mamoda2.5: Enhancing Unified Multimodal Model with DiT-MoE (2026) Please give a thumbs up to this comment if you found…