Paper page - PixelEyes: Decoupling Perception and Reasoning for Pinpoint Visual Evidence Seeking
…a new benchmark with expert-resynthesized data. Generated by Qwen/Qwen2.5-Coder-32B-Instruct This paper explores multi-turn visual reasoning and observes that MLLMs repeatedly fail to localize the target…