cd /news/artificial-intelligence/terminal-agents-a-survey-of-ai-agent… · home topics artificial-intelligence article
[ARTICLE · art-108236] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Terminal Agents: A Survey of AI Agents in Command-Line Environments

A new survey from arXiv (2608.20485v1) defines terminal agents as AI systems whose action-observation loop is mediated by command-line execution, and proposes a seven-dimensional competence profile to unify research across software engineering and other domains. The survey finds that agent behavior is shaped by model, interface, harness, runtime, and environment, and that current evaluations overemphasize final outcomes while underreporting process quality, recovery, and governance. It calls for explicit reporting of system and runtime conditions with replayable traces to improve comparability.

read1 min views1 publishedAug 24, 2026

arXiv:2608.20485v1 Announce Type: new Abstract: Large language model agents increasingly act through terminals, yet existing surveys disperse terminal-mediated behavior across software engineering, tool use, and computer-use research. We regard terminal agents as systems whose dominant progress-bearing action--observation loop is mediated by terminal command execution, textual feedback, and stateful environment interaction. Using terminal-mediated execution as an organizing lens, this survey establishes workload-level boundaries and connects system architecture, competence acquisition, and evaluation through a seven-dimensional terminal competence profile. Our synthesis shows that realized behavior is jointly shaped by the model, interface, harness, runtime, and environment. Executable trajectories ground learning in action consequences, verification, and recovery, whereas prevailing evaluations emphasize final outcomes and expose process quality, recovery, and governance unevenly. Bounded fixed-condition diagnostics illustrate two implications: benchmark families expose different process signals, and matched system comparisons reveal benchmark-dependent performance and limits of component attribution. These findings motivate explicit reporting of system and runtime conditions, supported by replayable traces and process-level evidence. The framework provides a unified basis for studying terminal-mediated agency across software engineering and emerging application domains.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/terminal-agents-a-su…] indexed:0 read:1min 2026-08-24 ·