2
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
The paper introduces Long-Horizon-Terminal-Bench, a benchmark designed to evaluate AI agents on complex, long-duration tasks using dense, reward-based grading.
Impact
42/100
Current rank score
1.8
Source tier
Tier 1
Category
Research
Firefly links to the original publisher. The summary above is AI-generated for orientation and may differ from the source. The “current rank score” decays over time so newer significant stories surface first.