AI News selected for Professionals and Decision Makers
Primary Research Stream

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading

06:00 · July 13, 2026 · arXiv cs.AI RSS

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading

AI agents have become capable of autonomously completing short, well-specified tasks. However, existing terminal benchmarks largely focus on simple problems that finish within minutes and are evaluated only by their final outcome. This setup overlooks intermediate progress and partial solutions, yielding sparse reward signals and an incomplete picture of agent capability. We introduce Long-Horizon-Terminal-Bench, a terminal benchmark of 46 long-horizon tasks spanning nine categories, including experiment reproduction, software engineering, multimodal analysis, interactive games, and scientific computing. Each task follows a Terminal-Bench-style setup with a reference solution or simulation engine, but is further decomposed into fine-grained graded subtasks. This design enables dense intermediate rewards and partial credit, allowing evaluation to capture not only whether an agent reaches the final goal, but also how far it progresses on open-ended workflows. Tasks in Long-Horizon-Terminal-Bench typically require hundreds of episodes and minutes to hours of execution, stressing long-horizon planning, long-context management, and iterative debugging rather than one-shot problem solving. We evaluate 15 frontier models and find that agents consume on average 9.9M tokens per task, with roughly 231 episodes and 85.3 minutes of execution time per run, making Long-Horizon-Terminal-Bench more demanding than prior terminal-based benchmarks. Even the strongest tested model achieves 15.2% pass@1 at a partial-reward threshold of 0.95 and 10.9% at a perfect-reward threshold of 1.0, while the mean pass rate across models is 4.3% and 1.7% under the two thresholds, respectively. These results reveal headroom for improvement. We further analyze failure modes and error patterns, and release Long-Horizon-Terminal-Bench to support future progress on long-horizon terminal agents.

Summary

Long-Horizon-Terminal-Bench addresses a gap in current terminal-agent evaluation by providing 46 tasks that require sustained execution across many steps rather than short, self-contained actions. The benchmark covers nine categories such as experiment reproduction, software engineering, multimodal analysis, interactive games, and scientific computing. Each task runs inside a containerized terminal environment with a natural-language goal, relevant files and tools, and a reference solution or simulator. Unlike prior terminal benchmarks that grade only the final state, Long-Horizon-Terminal-Bench decomposes every task into fine-grained subtasks that receive partial credit, producing dense reward signals throughout long trajectories.

Tasks are designed to last tens of minutes to hours and typically demand hundreds of episodes. Across evaluations of 15 frontier models, agents averaged 9.9 million tokens, 231 episodes, and 85.3 minutes of wall-clock time per task under a 90-minute timeout—roughly an order of magnitude more demanding than earlier suites such as Terminal-Bench 2. The strongest model reached a 15.2 percent pass rate at a 0.95 partial-reward threshold and 10.9 percent at a perfect 1.0 threshold; mean performance across all models was 4.3 percent and 1.7 percent under the same thresholds. These figures indicate that current agents struggle to maintain coherent plans, manage accumulating context, and recover from errors over extended horizons.

The authors also examine recurring failure patterns, including premature termination, weak self-verification, and inability to complete the final stages even after substantial intermediate progress. By releasing the benchmark and its evaluation harness, the work supplies a concrete testbed for measuring and improving the long-horizon reliability of terminal agents.

Why it matters

This research is highly relevant for Dutch AI researchers and developers focusing on autonomous agents and LLM evaluation. The introduction of dense reward grading provides a more nuanced framework for testing agent robustness, which is critical for deploying reliable AI systems in Dutch enterprises.

More in this beat
ai-agentsevaluation-benchmarksexperimental-benchmarksllm-agentslong-horizon-terminal-benchnovel-methodologiestechnical-rigor
ARCANA: A Reflective Multi-Agent Program Synthesis Framework for ARC-AGI-2 Reasoning

06:00 · July 13, 2026

ARCANA: A Reflective Multi-Agent Program Synthesis Framework for ARC-AGI-2 Reasoning

This highly technical paper is directly relevant to AI researchers and advanced practitioners in the Netherlands working on AGI, multi-agent systems, and abstract reasoning. Its focus on achieving state-of-the-art results under strict hardware constraints makes it highly actionable for Dutch research labs and AI-driven SMEs looking to deploy efficient reasoning models.

Relevance 85 · Audience 95

Controlling Tool Use with Heading-Specific Activation Steering

06:00 · July 8, 2026

Controlling Tool Use with Heading-Specific Activation Steering

This research provides advanced techniques for controlling LLM agent behavior, which is crucial for Dutch AI researchers developing reliable and efficient AI systems. Understanding and steering tool use aligns with the EU's push for transparent and predictable AI deployments.

Relevance 85 · Audience 95

What Drives Interactive Improvement from Feedback?

06:00 · July 1, 2026

What Drives Interactive Improvement from Feedback?

This research is highly relevant for Dutch AI researchers and developers building LLM agents, as it provides a rigorous framework to evaluate feedback mechanisms. It aligns with the EU's push for robust, transparent AI by highlighting the need for proper baselines (repeated attempts) rather than misleading multi-turn accuracy metrics.

Relevance 85 · Audience 95

SkillHarness: Harnessing Safe Skills for Computer-Use Agents

06:00 · June 23, 2026

SkillHarness: Harnessing Safe Skills for Computer-Use Agents

This research directly supports the Dutch and EU strategic focus on safe, ethical, and reliable AI deployment. For researchers and advanced practitioners in the Netherlands, it provides actionable methodologies to build autonomous agents that comply with stringent safety constraints in dynamic environments.

Relevance 85 · Audience 90

Is it agentic enough? Benchmarking open models on your own tooling

02:00 · June 18, 2026

Is it agentic enough? Benchmarking open models on your own tooling

It provides ML Engineers with actionable insights and a new open-source tool to benchmark and optimize their own libraries for agentic use. Understanding the trade-offs in token consumption and latency across different model sizes is crucial for building cost-effective and reliable AI systems.

Relevance 85 · Audience 95

FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive Fraud

06:00 · August 20, 2026

FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive Fraud

This research is highly relevant for Dutch AI researchers and the strong local fintech and banking sector exploring customer-facing LLM agents. It provides a rigorous, reproducible framework to test agent compliance and security against fraud, aligning with strict EU financial and AI regulations.

Relevance 85 · Audience 95

Measuring Cross-Task Behavioral Consistency in Language Model Agents

06:00 · August 17, 2026

Measuring Cross-Task Behavioral Consistency in Language Model Agents

The article provides a novel, quantifiable method for assessing the reliability and behavioral consistency of AI agents, which is crucial for compliance with EU AI regulations and the Dutch focus on transparent AI. Researchers can directly apply the open-source BCM framework to evaluate and improve the predictability of enterprise AI deployments.

Relevance 85 · Audience 95