Paper page - MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations
… The following papers were recommended by the Semantic Scholar API CEO-Bench: Can Agents Play the Long Game? …
… The following papers were recommended by the Semantic Scholar API CEO-Bench: Can Agents Play the Long Game? …
… In this paper, we propose Ctx2Skill, a self-evolving framework that autonomously discovers, refines, and selects context-specific skills without human supervision or external feedback. …
… Published on Jun 1 Submitted by Jiaheng Liu on Jun 4 NJU-LINK Lab Authors: , , , , , , , , , , , , Abstract MMG2Skill framework converts web-based procedural guides into executable skills through closed-loop learning, improving agent performance across GUI control, gameplay, and card play tasks. …
… Further analysis reveals that while agents often implement recognizable mechanics, they struggle to deliver complete games with sufficient content, functional visual feedback, and coherent presentation. See https://tongxuluo.github.io/ gamecraft-bench -website for demos, code, and data. …
… We propose CapCode, a framework for constructing coding datasets with randomized tests whose best achievable non-cheating performance is deliberately capped below one. …
… Each task sustains at least 12 hours of continuous agent operation under rich, multilevel feedback , and is built through substantial expert effort. …
… A researcher agent R run as a coding agent reads the inner-loop source code, edits system prompts, feedback functions, helper libraries, and iteration logic, runs evaluations, and decides what to keep, following the autoresearch paradigm. …
… The following papers were recommended by the Semantic Scholar API SPIKE: An Adaptive Dual Controller Framework for Cost-Efficient Long-Horizon Game Agents 2026 Exploring Cross-Scenario Generality of Agentic Memory Systems: Diagnostics and a Strong Baseline 2026 LLM Agents Are Latent Context Manager… …
Papers arxiv:2605.28897 Review Arcade: On the Human Alignment and Gameability of LLM Reviews Published on May 27 Submitted by Jan Strich on Jun 2 Hub of Computing and Data Science HCDS - G4KMU Authors: , , Jan Strich Abstract Empirical analysis reveals limited alignment between LLM-generated review… …
… View arXiv page View PDF Add to collection Community Real-world agent learning is often constrained by costly environment interactions, such as running time-consuming experiments or obtaining human feedback. …