Hi HF team,
I’d like to register IntelligenceLab/Long-Horizon-Terminal-Bench as a Benchmark so its leaderboard can aggregate community evaluation results across the Hub.
Dataset: [IntelligenceLab/Long-Horizon-Terminal-Bench · Datasets at Hugging Face](https://huggingface.co/datasets/IntelligenceLab/Long-Horizon-Terminal-Bench)
Paper: [[2607.08964] Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading](https://arxiv.org/abs/2607.08964)
LHTB is a 46-task benchmark testing how well LLM agents sustain useful work in a containerized terminal over hundreds of steps, using hidden, rebuild-from-artifact verifiers rather than self-reported progress. Tasks run via the Harbor / Terminal-Bench 2.0 harness.
eval.yaml is already in the repo root (evaluation_framework: harbor) and validated on push.
Could you add this to the Benchmark allow-list? Thanks!