{"slug": "add-intelligencelab-long-horizon-terminal-bench-to-the-benchmark-allow-list", "title": "Add IntelligenceLab/Long-Horizon-Terminal-Bench to the Benchmark allow-list", "summary": "IntelligenceLab requests Hugging Face to add Long-Horizon-Terminal-Bench (LHTB) to the Benchmark allow-list. LHTB is a 46-task benchmark that tests LLM agents on long-horizon terminal tasks using hidden verifiers, with an eval.yaml already in the repository root.", "body_md": "Hi HF team,\n\nI’d like to register IntelligenceLab/Long-Horizon-Terminal-Bench as a Benchmark so its leaderboard can aggregate community evaluation results across the Hub.\n\nDataset: [IntelligenceLab/Long-Horizon-Terminal-Bench · Datasets at Hugging Face](https://huggingface.co/datasets/IntelligenceLab/Long-Horizon-Terminal-Bench)\n\nPaper: [[2607.08964] Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading](https://arxiv.org/abs/2607.08964)\n\nLHTB is a 46-task benchmark testing how well LLM agents sustain useful work in a containerized terminal over hundreds of steps, using hidden, rebuild-from-artifact verifiers rather than self-reported progress. Tasks run via the Harbor / Terminal-Bench 2.0 harness.\n\neval.yaml is already in the repo root (evaluation_framework: harbor) and validated on push.\n\nCould you add this to the Benchmark allow-list? Thanks!", "url": "https://wpnews.pro/news/add-intelligencelab-long-horizon-terminal-bench-to-the-benchmark-allow-list", "canonical_source": "https://discuss.huggingface.co/t/add-intelligencelab-long-horizon-terminal-bench-to-the-benchmark-allow-list/178048#post_1", "published_at": "2026-07-20 18:59:46+00:00", "updated_at": "2026-07-20 19:10:46.626129+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-agents", "ai-research", "ai-tools"], "entities": ["IntelligenceLab", "Hugging Face", "Long-Horizon-Terminal-Bench", "Harbor", "Terminal-Bench 2.0"], "alternates": {"html": "https://wpnews.pro/news/add-intelligencelab-long-horizon-terminal-bench-to-the-benchmark-allow-list", "markdown": "https://wpnews.pro/news/add-intelligencelab-long-horizon-terminal-bench-to-the-benchmark-allow-list.md", "text": "https://wpnews.pro/news/add-intelligencelab-long-horizon-terminal-bench-to-the-benchmark-allow-list.txt", "jsonld": "https://wpnews.pro/news/add-intelligencelab-long-horizon-terminal-bench-to-the-benchmark-allow-list.jsonld"}}