Show HN: System One Harness (SOH), the harness for System One models HarnessRouter released System One Harness (SOH), an open-source agent loop that turns a System One decision model into an agent by compiling an environment's finite action space into typed questions and gating each decision by confidence. The first supported model is Jev by TypeSafe, available through OpenRouter or TypeSafe directly, and a built-in order fulfilment demo completed in 5 steps with a wall time of 0.98s and a cost of $0.000208. The harness supports pluggable environments including Python processes, MCP servers, and web pages, and serves the loop through the Unified Harness Protocol for streaming, continuation, cancellation, and discovery. Turn a System One decision model into an agent loop. System One Harness observes an environment, compiles its finite action space into typed questions, gates each decision by confidence, executes the chosen action, and records the complete trace. One model call per step. No generated actions. A probability on every transition. Jev playing a live browser game through System One Harness — one typed decision per step, with no generated control text. The first supported model is Jev https://typesafe.ai by TypeSafe, available through OpenRouter or TypeSafe directly. Tip Start here: Run the example quickstart · Understand the loop how-it-works · Connect an environment connect-an-environment · Read the design https://github.com/HarnessRouter/SystemOneHarness/blob/main/docs/design.md git clone https://github.com/HarnessRouter/SystemOneHarness.git cd SystemOneHarness pip install -e . export OPENROUTER API KEY=sk-or-... or TYPESAFE API KEY=... s1 run --env order:ship fastest gift This runs the built-in order fulfilment environment against the live model: goal: Order B-220 is a gift: note it, then ship it by the fastest carrier. model: ~typesafe/jev-latest 0 add note note='gift' p=0.96 241 ms 1 pick item item='scarf' p=1.00 269 ms 2 pack p=0.99 166 ms 3 choose carrier carrier='express' p=0.98 151 ms 4 ship p=0.93 152 ms status=completed reason=environment terminal steps=5 wall=0.98s cost=$0.000208 Each step shows the selected action, its weakest required probability, and the model round trip. The final line records how the run ended, how long it took, and what it cost. | Finite actions | The model chooses only from actions and parameter values declared by the environment. | | Confidence gates | Read, write, and destructive actions can require different probability thresholds. | | Explicit outcomes | Every run ends as completed, incomplete, failed, or cancelled with a structured reason. | | Complete traces | State, questions, distributions, verdicts, results, latency, and usage are recorded step by step. | | Pluggable environments | Drive Python processes, MCP servers, or web pages with the same controller. | | UHP compatibility | Serve the loop through the Unified Harness Protocol for streaming, continuation, cancellation, and discovery. | Games, live feeds, and other moving environments can return "realtime": true . The controller then treats a refused or repeated decision as a clock tick, keeps the model's history short, and lets the last action remain active until it changes. ┌──────────────────────────────────────────────────────────┐ │ controller │ goal ────► │ observe ─► compile ─► encode ─► decide ─► gate ─► execute │ ────► trace │ ▲ │ │ │ └────────────── environment ◄───────────────┘ │ └──────────────────────────────────────────────────────────┘ 1. Observe. The environment reports text, structured fields, candidates, and terminal state. 2. Compile. The available actions become typed choice , noul , and score questions. 3. Encode. Goal, observation, bounded history, and memory become a state within the model budget. 4. Decide. The model answers the action, its parameters, and the goal check in one request. 5. Gate. The weakest required probability must clear the selected action's risk threshold. 6. Execute. The environment applies the action and returns the next state. finish and escalate are actions, not generated prose. The controller always knows why it stopped. Read the measured design and architecture → https://github.com/HarnessRouter/SystemOneHarness/blob/main/docs/design.md Choose the smallest boundary that fits your system. | Environment | Use it when | Start with | |---|---|---| | Action space + process | You own a local program or service loop. | s1 run --actions actions.yaml --env-cmd "python3 env.py" --goal "..." | | MCP server | Your tools already expose enumerable inputs over MCP. | s1 run --mcp "python -m your server" --goal "..." | | Browser | The task is expressed through DOM controls in Chrome. | s1 run --browser --headless --start-url https://example.com --goal "..." | | Python | You want an in-process integration. | Subclass Environment and implement observe and execute . | An action space is YAML or the same structure in Python. Every parameter must be enumerable. instructions: - Move the order to shipped, or cancel it when the goal says so. actions: choose carrier: description: Select a carrier for the packed order. risk: write params: carrier: from: available carriers ship: description: Hand the packed order to the selected carrier. risk: destructive gate: read: 0.5 write: 0.6 destructive: 0.8 finish: 0.5 Parameters can use fixed choices , observation candidates , a boolean flag , or ordered levels . Free text is rejected because a System One model does not generate text. The harness lists an MCP server's tools, compiles supported schemas into actions, and explains every unsupported tool instead of silently dropping it. pip install -e ". mcp " s1 tools --mcp "python -m systemone harness.envs.order mcp" s1 run --mcp "python -m systemone harness.envs.order mcp --scenario ship fastest gift" \ --goal "Order B-220 is a gift. Ship it by the fastest carrier." An observe tool provides state. An optional reset tool starts a run. Every other compatible tool becomes an action. MCP annotations determine whether the action is read, write, or destructive. The browser environment uses Browser Use https://github.com/browser-use/browser-use to turn visible DOM controls into a finite action space. Text comes from named values supplied by the caller. The model chooses values by name and never writes them. pip install -e ". browser " s1 run --browser --headless --start-url https://example.com/book \ --text name=Customer --text email=user@example.com \ --goal "Book a table at 19:30 with a window seat." Browser setup, measurements, and limits → https://github.com/HarnessRouter/SystemOneHarness/blob/main/docs/browser-use.md Expose any configured loop as a Unified Harness Protocol https://unifiedharnessprotocol.org server: export OPENROUTER API KEY=sk-or-... s1 serve --api-key choose-a-secret --port 8710 input → goal function call → selected action function call output → environment result reasoning → distribution and gate verdict previous response id → continued environment and history Streaming emits each item as it happens. Cancellation lets the current model step finish and records it. The included report passes all 40 checks in the UHP core conformance class. Five live runs per scenario on 2026-09-19 with typesafe/jev-1.13-20260917 through OpenRouter: | Scenario | Goal met | Mean steps | Mean model latency | Mean wall time | Cost per run | |---|---|---|---|---|---| | Ship by cheapest carrier | 5/5 | 6.0 | 241 ms | 1.45 s | $0.000265 | | Ship fastest and add gift note | 5/5 | 5.0 | 199 ms | 0.99 s | $0.000214 | | Cancel a fraudulent order | 5/5 | 1.0 | 197 ms | 0.20 s | $0.000044 | The benchmark proves the controller, compiler, gate, and model can complete these small deterministic tasks. It does not claim the same result for ambiguous state, arithmetic, dates, or long irrelevant context. Inspect the raw benchmark rows → https://github.com/HarnessRouter/SystemOneHarness/blob/main/docs/reports/bench-2026-09-19.json s1 run Run one goal against a built-in, process, MCP, or browser environment s1 tools Inspect how an MCP server compiles into supported actions s1 serve Expose a configured loop as a UHP server s1 bench Run the built-in live benchmark Use s1