Paper page - K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts
…Generated by Qwen/Qwen2.5-Coder-32B-Instruct Frontier model evaluations are shifting from foundational capabilities (e.g., instruction following and reasoning) toward compositional, agentic ones, but Korean agentic benchmarks remain scarce…