Paper page - Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models
…Generated by Qwen/Qwen2.5-Coder-32B-Instruct Vision-language models (VLMs) project images into hundreds to thousands of visual tokens , making decoder inference expensive in both attention computation and KV-cache…