Paper page - LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks
…https://lh-harness.pages.dev Interesting approach! Separating task state from execution and verifying it independently seems like a practical way to improve reliability for long-horizon AI agents. The benchmark improvements…