Paper page - WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
…We release the tasks, code, and containerized tooling to support reproducible evaluation. github repo: https://github.com/InternLM/WildClawBench leaderboard: https://internlm.github.io/WildClawBench/ This is an automated message from the…