FrontierAgent Runs Its Terminal Agent on the Same Engine as Its Benchmarks The open-source FrontierAgent project ships a terminal agent runtime whose workflow engine is the same one powering the benchmark runner used to evaluate Apodex models, according to its README. The runtime separates the generic agent loop, tool plugins, ReAct and Agent Team workflows, the terminal CLI/TUI, and a public benchmark harness into distinct directories, and runs agents in a task-scoped filesystem with read-only /inputs, a /workspace, and persistent /outputs. Mutating operations require approval unless --yes is set, sessions are checkpointed and traced locally, and Agent Team parallelism adds to benchmark concurrency, which the README flags as a load warning. My reading of the FrontierAgent README is that its interesting claim is structural rather than a feature list: the workflow engine behind the frontier-agent terminal app is, according to the README, the same one that powers the benchmark runner used to evaluate Apodex models. That makes the terminal product feel closer to a published test harness. If you want to see how these agent workflows behave on long-horizon research and file work, you can run them yourself, on your own tasks, against your own OpenAI-compatible endpoint. The README describes the project as an open-source agent runtime, terminal product, and evaluation suite. It also says the framework, tools, workflows, and evaluation layer remain separate so each can be reused independently. The repository layout reflects that with distinct directories: frontier agent/ for the generic loop, scheduling, registries, AgentBus, and observers; plugins/tools/ for tool implementations; workflows/ for the ReAct and Agent Team pipelines; apodex/ for the terminal CLI and TUI; and benchmarks/ for the public harness plus bundled FrontierSearchBench and FrontierChallenge. The TUI ships two native workflows, and the README scopes them differently. In react mode, one stateful agent researches, reads files, writes deliverables, runs commands, and iterates in a task-scoped sandbox. The README pitches it for focused research, repository analysis, and document or file work, using the tui workflow profile. In agent team mode, a coordinator maintains a task board, delegates independent work to parallel sub-agents, collects their reports, and synthesizes the result. The README adds that the coordinator dispatches bounded parallel assignments, receives structured reports, and can use an optional fast reporter for final evidence review. The task board is visible: add task and update task events show up live in the TUI sidebar with pending, active, completed, blocked, and cancelled states. Picking a mode is a flag: uv run frontier-agent --mode react --cwd /path/to/project uv run frontier-agent --mode agent team --cwd /path/to/project One behavior differs between the modes in a way that matters mid-run. You can type while an agent works, and the instruction is queued and injected at the next safe turn boundary without discarding the active run. In Agent Team mode, that input steers the coordinator, while sub-agents that are already running are allowed to finish. The part I would read first as an operator is the filesystem. Shell and file tools share one task-scoped filesystem: /inputs is read-only, /workspace holds working state, and /outputs holds persistent deliverables. The README states that authorization and sandbox failures are fail-closed. On macOS and Docker, /outputs maps to .apodex/runs/