Paper page - PACE: A Proxy for Agentic Capability Evaluation
…predicts expensive agentic LLM benchmark performance using a small subset of atomic evaluation instances, achieving high accuracy at a fraction of the cost. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Evaluating…